3 ms·
On a side note, isn't the choice of exactly four unrolls very architecture specific? As in, it works, but may be sub-optimal for your specific machine. I've don
by csl 10y ago
On a side note, isn't the choice of exactly four unrolls very architecture specific? As in, it works, but may be sub-optimal for your specific machine. I've done the exact same thing myself, and IIRC its performance varied a lot between which ISA it was compiled for.
This is almost what Duff's device solves, except then you need to know the length beforehand.
- rikkus 10y agoAbsolutely. It's a (possible) optimisation that is either based on evidence (seems likely, because DJB) or hope. Actual behaviour is impossible to predict on untested platforms. My assumption is that DJB tested this locally and found enough of a speedup that it was worth it, considering the very low added complexity and risk of major degradation / defects on untested platforms.