3 ms·
The first C result is absurd, not sure how the author could have gotten it. First of all, the code as written will just optimize to nothing, so we need to add
by devit 11y ago
The first C result is absurd, not sure how the author could have gotten it.
First of all, the code as written will just optimize to nothing, so we need to add an asm("" : "=g" (s) : "0" (s)) in the loop to stop strength reduction and autovectorization, and we need to return the final value to stop dead code elimination.
Once that is done, the result is more than 2 billion iterations per second on a ~3 GHz Intel desktop CPU, while the author gives an absurd value of 500m iterations which could not have been possibly obtained with any recent Intel Xeon/Core i5/i7 CPU.
BTW, the assembly code produced is this:
1:
add $0x1,%edx
add $0x1,%esi
cmp %eax,%edx
jne 1b
Which is unlikely to take more than 1/2 cycles to execute on any reasonable CPU as my test data in fact shows.
- CydeWeys 11y agoWell there's always flags to prevent compiler optimizations, or maybe the example was purposefully presented in readable C, not whatever hack you'd need to do to bypass optimization. Inline assembly isn't exactly C anymore. But yeah, I was surprised by the number of operations per second too. I was thinking it had to be over a billion.