3 ms·
> mov instructions from register to register make up for more than 60% of the time spent in the critical section of the code, while we would expect most of the
by CUViper 11y ago
> mov instructions from register to register make up for more than 60% of the time spent in the critical section of the code, while we would expect most of the time to be spent xoring and anding. I have not investigated why this is the case, ideas welcome
If you're not using precise events, then the instruction addresses reported by perf will have some skid. This is a small cpu delay from when a performance counter overflows to when the interrupt actually freezes state.
You can choose precise sampling for some events, depending on the CPU. Try "-e cycles:pp" for instance.
0,09 │ mov (%rax,%r8,4),%eax
29,32 │ mov %r14,%r8
I think this first mov from memory is likely to be your true cycle eater, much more so than the second mov reg-reg or any single xor/and operations. But don't optimize based on my hunch - measure it precisely first! If memory access proves to be your slowdown, then you can try optimizing your access patterns.
- wyager 11y agoAlso, because of how complicated modern pipelining is, some instructions that you wouldn't expect to take a long time do take a long time (usually because they're waiting on a mov from memory to finish). In this case, the mov from memory could be throwing up a hazard that's blocking the mov from register.