6 ms·
Tight loops of SIMD operations seems like something that might be more convenient to just implement directly in assembly? So you don't need to babysit the compi
by jonathrg 2y ago
Tight loops of SIMD operations seems like something that might be more convenient to just implement directly in assembly? So you don't need to babysit the compiler like this.
- adgjlsfhk1 2y agoCounterpoint: that makes you rewrite it for every architecture and every datatype. A high level language makes it a lot easier to get something more readable, runable everywhere, and datatype generic.
- brrrrrm 2y agoYou almost always have to for good perf on non-trivial operations
- menaerus 2y agoIf you're ready to settle on much lower IPC then yes.
- adgjlsfhk1 2y agoThe Julia version of this (https://github.com/miguelraz/StagedFilters.jl https://github.com/miguelraz/StagedFilters.jl) is very clearly generating pretty great code. For `smooth!(SavitzkyGolayFilter{4,3}, rand(Float64, 1000), rand(Float64, 1000))`, 36 out of 74 of the inner loop instructions are vectorized fmadds (of one kind or another). There are a few register spills that seem plausibly unnecessary, and some dependency chains that I think are an inherent part of the algorithm, but I'm pretty that there isn't an additional 2x speed to get here.
- menaerus 2y agoI can write SIMD code that's 2x the speed of the non-vectorized code but I can also rewrite that same SIMD code so that 2x becomes 6x. Point being, not only that you can get 2x on top of initial 2x SIMD implementation but usually much more. Whether or not you see SIMD in the codegen is not a testament of how good the implementation really is. IPC is the relevant measure here.
- adgjlsfhk1 2y agoIPC looks like 3.5 instructions per clock. (and the speed for bigger inputs will be memory bound rather than compute bound).
- menaerus 2y agoIf true this would be between ok-ish and decent result but since you didn't say anything about the experiment or how you ran it or anything about the HW it's hard to say more. I assume you didn't use https://github.com/miguelraz/StagedFilters.jl/blob/master/test/benchmarks.jl https://github.com/miguelraz/StagedFilters.jl/blob/master/te... since that benchmark is pretty much flawed for more than a few reasons: - Operates with the dataset that cannot fit even in the L3 cache on most consumer machines: ~76MB (float64) and ~38MB (float32) - Intermixes python with the julia code - Not even a single repetition - no statistical significance - Does not show what happens with different data distributions, e.g. random, but only uses a monotonically increasing sequence
- ChrisRackauckas 2y agoDid you link the wrong script? The script you show runs everything to statistical significance using Chairmarks.@b. Also, I don't understand what the issue would be with mixing Python and Julia code in the benchmark script. The Julia side JIT compiles the invocations which we've seen removes pretty much all non-Python overhead and actually makes the resulting SciPy calls faster than doing so from Python itself in many cases, see for example https://github.com/SciML/SciPyDiffEq.jl?tab=readme-ov-file#measuring-overhead https://github.com/SciML/SciPyDiffEq.jl?tab=readme-ov-file#m... where invocations from Julia very handily outperform SciPy+Numba from the Python REPL. Of course, that is a higher order function so it's benefiting from the Julia JIT in other ways as well, but the point is in previous benchmarks we've seen the overhead floor as so low (~100ns IIRC) that it didn't effect benchmarks negatively for Python and actually improved many Python timings in practice. Though it would be good to show in this case what exactly the difference is in order to isolate any potential issue, I would be surprised if it's more than a 100ns overhead for an invocation like this and with 58ms being the benchmark size, that cost is well below the noise floor). Though trying different datasets is of course valid. There's no reason to reject a benchmark just because it doesn't fit into L3 cache, there's many use cases for that. But it does not mean that all use cases are likely to see such a result.
- jonathrg 2y agoFor libraries, yes that is a concern. In practice you're often dealing with one or a handful of instances of the problem at hand.
- anonymoushn 2y agoI'm happy to use compilers for register allocation, even though they aren't particularly good at it
- menaerus 2y agoDefinitely. And I also found the title quite misleading. It's the auto-vectorization that this presentation is trying to cover. Compile-time SIMD OTOH would mean something totally different, e.g. computation during the compile-time using the constexpr SIMD intrinsics. I'd also add that it's not only about babysitting the compiler but you're also leaving a lot of performance off the table. Auto-vectorized code, generally speaking, unfortunately cannot beat the manually vectorized code (either through intrinsics or asm).