3 ms·
I wonder if this code could be even further speed up, by taking into account the througput versus the latency of the FMA operator. If the latency is 4 cycles an
by guyomes 3y ago
I wonder if this code could be even further speed up, by taking into account the througput versus the latency of the FMA operator. If the latency is 4 cycles and the througput is 0.5 cycle for example, that means that doing n dependent operation will cost 4n cycles, and n independent operations will cost 0.5 cycles, that is an 8x speed up. By dependent operation, I mean that the input of an operation depends on the output the previous operation.
In the current final code, each loop has 3 independent FMAs. And in each loop iteration, the 3 FMAs are dependent on the FMAs from the previous iteration.
So one thing that could speed up the computation is to do 6 independent FMAs per loop iteration instead of 3. This can be done as follow. Instead of doing cumulative sums from 0 to n-1, one could split the computation of the cumulative sums from 0 to n/2-1 and from n/2 to n-1. So the main loop goes only from 0 to n/2-1, and at the k-th loop iteration, we do the 2x3 FMAs corresponding the entries of indices k and n/2+k. And at the end we add the cumulative sums.
- ashvardanian 3y agoI've had a similar hypothesis, but it didn't work out for me, only bloating the codebase. Please lmk if you have a different outcome :)
- anonymoushn 3y agoI don't know about this code in particular but if you're mostly doing math you tend to run out of architectural registers before you achieve the full throughput of the instructions you care about :(
- guyomes 3y agoThat makes sense. I did observe significant speed up using this technic for evaluating a polynomial on several values. I used Hörner algorithm and indeed the number of registers is very small in this case. If the lack of registers is really the bottleneck, a variant might to use both zmm registers for some cumulative sums and ymm registers for some others. In this case the speed up might be less spectacular though. Edit: actually, I just discovered that zmm registers overlap the ymm registers, so the only registers left are the ones from the FPU.
- gpderetta 3y agoAVX512 added 16 more vector registers. Isn't 32 registers enough to unroll by 8?