4 ms·
This is probably a bad idea and I would bet (and probably win) that less aggressive unrolling would be better in pretty much all cases except a micro-benchmark.
by pslam 12y ago
This is probably a bad idea and I would bet (and probably win) that less aggressive unrolling would be better in pretty much all cases except a micro-benchmark.
This kind of trick worked well on CPUs in the 1980s (as jvoorhis points out in another comment). It works wonder for micro-benchmarks. It is a net loss for any CPU and workload where instruction cache miss rate can severely impact performance. That is to say, pretty much everything made after 1990s and doing more than just shuffling memory around.
The problems are: 1) this is a large fraction of I-Cache size, 2) unrolling doesn't actually save much (or sometimes anything) on out-of-order multi-issue cores.
First off: extremely aggressive unrolling results in increasing I-cache thrashing. This is usually not apparent in micro-benchmarks, because the entire working code set fits in the I-cache. It becomes apparent if the unrolling is stupid enough to be larger than I-cache (or usually just close enough to it), because then you're spending just as much (usually more) time re-fetching instructions from memory as you are moving data around. That's why picking the best result from a micro-benchmark is misleading. In real use cases, you find that functions like this one tend to cause a "global slow down", with no specific function accounting for it, because they're greatly reducing the effectiveness of the I-cache.
Also, unrolling doesn't really do much these days. You can do better than REP MOVSD and friends - go read Intel's documentation for their recommended sequences (which has a nasty habit of changing every damn CPU generation). Modern out-of-order, multi-issue cores can very easily saturate their load/store units, while also executing all those pesky loop counters and branches. In fact, that's the whole point of being out-of-order and multi-issue: keeping units busy.
So, please can everybody stop using Duff's device, unless: 1) you're targeting a CPU which has no I-cache (e.g tightly coupled ROM/RAMs), 2) you really do have a working set where this fits well enough to be better, 3) you're trying to cheat on benchmarks, or 4) you're trying to obfuscate your binaries against reverse engineering.
- eloff 12y agoRep movsd is only recommended for large clears/copies by Intel. Indeed the commit message states: REP MOVSQ and REP STOSQ have a really high startup overhead. Use a Duff's device to do the repetition instead.
- deleted 12y ago[deleted]
- mc_hammer 12y agothanks for the post - ive never read any of this stuff before. ive only been a dev for 18 years ¬_¬
- chrisbennet 12y agoKids these days.. ;-)