3 ms·
Thanks for this clear explanation! Regarding your example with 1000 muls followed by 1000 loads... That's why in my experiments I interleaved loads and bswaps,
by dendibakh 9y ago
Thanks for this clear explanation!
Regarding your example with 1000 muls followed by 1000 loads... That's why in my experiments I interleaved loads and bswaps, because that's what (hopefully) every decent compiler will do.
- nkurz 9y agoI interleaved loads and bswaps, because that's what (hopefully) every decent compiler will do Your optimism about compiler behavior is charming. I think you may be disappointed if you expect compilers to interleave loads just because this approach works better on modern processors. My experience has been that GCC goes out of its way to hoist all of your carefully interleaved loads to a big block at the top. I'm guessing it does this because it's a simple heuristic that was often helpful before out-of-order processors became common. Normally, the performance impact of this is very small, but when it does exist, it's usually negative. If I recall, ICC does a better job of interleaving, or at least leaving things interleaved.
- dendibakh 9y agoWell, yeah. I don't know much about the current state of the art (because I don't touch the CodeGen on a daily basis), but I kind of look into the future with hope that compilers will better handle at least those "simple" cases. :)
- BeeOnRope 9y agoCompilers are rarely going to transform the 1000/1000 example into the 10/10 one, even if they were smarter. Often such a transformation is simply impossible: the effect of interleaving the instructions may be different than what the source dictates. Also the 1000/1000 example probably doesn't arise as a long stream of explicit instructions: it is probably just a short loop with 1000 iterations! That makes it even less likely that the compiler will simply start interleaving various instructions following the loop into the loop body somehow.