5 ms·
That makes sense. I did observe significant speed up using this technic for evaluating a polynomial on several values. I used Hörner algorithm and indeed the nu
by guyomes 3y ago
That makes sense. I did observe significant speed up using this technic for evaluating a polynomial on several values. I used Hörner algorithm and indeed the number of registers is very small in this case.
If the lack of registers is really the bottleneck, a variant might to use both zmm registers for some cumulative sums and ymm registers for some others. In this case the speed up might be less spectacular though.
Edit: actually, I just discovered that zmm registers overlap the ymm registers, so the only registers left are the ones from the FPU.
- gpderetta 3y agoAVX512 added 16 more vector registers. Isn't 32 registers enough to unroll by 8?