4 ms·
This is just an educated guess, but the cmov is probably going to stall the execution stage of the pipeline, which could end up backing up the front-to-back-end
by bertr4nd 11y ago
This is just an educated guess, but the cmov is probably going to stall the execution stage of the pipeline, which could end up backing up the front-to-back-end queue, the reorder buffer, and perhaps even fetch itself. You're getting no help from the branch predictor at this point since the dependency runs through an ALU instruction, whereas in the branchy code, you at least have a coinflip chance of predicting the right direction.
- acqq 11y agoSee my answer to the other post. I still don't see what's going on. I'd expect that the access to the non-cached RAM dominates in big arrays, and we see that for short arrays CMOV is faster. There are tools to actually figure out what's going on, Intel can measure cache misses etc. But I'd like at least ASM codes and the example of indexes in one and another case, if they are very different that's the best explanation.
- bertr4nd 11y agoI think the other poster had it backwards. I'd expect CMOV to perform worse with high memory access latencies (which it does), because it stalls the pipeline. With low access latencies the pipe doesn't stall (for long) anyways, and you avoid the branch miss overhead.
- acqq 11y agoThanks, you motivated me to find this Linus' take about the CMOV stalls: http://yarchive.net/comp/linux/cmov.html http://yarchive.net/comp/linux/cmov.html Basically, if the direction is predictable, jump can be faster because the mov is then "unconditional." The strange thing is that the binary search on average shouldn't be predictable. So it's still the question what was measured there. Maybe always an element on the position a[0], even when the array was big?