3 ms·
It's a performance trade off. It's not really surprising that GCC would make different performance trade offs when optimizing for different microarchitectures.
by aij 9y ago
It's a performance trade off. It's not really surprising that GCC would make different performance trade offs when optimizing for different microarchitectures. (let alone most likely across different versions of GCC)
Why would you expect GCC to optimize for Pentium 4 (Netburst) in this day and age? (Especially given that the article is talking about Skylake.)
- raverbashing 9y agoI don't expect it to optimize to P4, are you talking about this? (it's one of the answers to that answer) > The quoted delay of 5 - 6 clocks is much better on later microarchitectures. For example from Sandy Bridge and Ivy Bridge, The Ivy Bridge inserts an extra μop only in the case where a high 8-bit register (AH, BH, CH, DH) has been modified And you see that even on later architectures the point of avoiding AH makes sense (which is the opposite of what that GCC code does)