3 ms·
I talked to the author of that article. He hasn't done testing on AMD processors but his guess was: micro-up fusion means the seven-instruction loop is actuall
by brucedawson 7y ago
I talked to the author of that article. He hasn't done testing on AMD processors but his guess was:
micro-up fusion means the seven-instruction loop is actually five micro-ops
Zen2 processors can retire five instructions per cycle
Therefore the loop runs at one iteration per cycle (wow!)
The cmp [r8] instruction occasionally has cache misses
This means that the seven instructions get synchronized such that the cmp [r8] instruction is the last of them to get retired in a seven-instruction block
Therefore the next instruction is usually the jne
TL;DR - the jne gets most of the samples because the cmp [r8] instruction is the most expensive.
- BeeOnRope 7y agoYeah I believe jne gets most of the samples because cmp [r8] is the most expensive, but there could be two separate reasons for that: Perhaps ETW shows you the precise instruction (i.e., "zero skid") that is slow to retire - this is not like a normal interrupt as described in the article but is available with some performance profiling events like 'cycles:ppp' on Linux perf (in particular, using the zero-skid PEBS events). In that case, the samples shows up on the jne, not the cmp likely because cmp/jne have fused, so basically get sampled as a single instruction and the samples show up pointing to the jump. The other scenario is that ETW shows you "skid 1" instructions, i.e., the instructions generally after the slow-to-retire ones (as described in the article), and cmp/jne didn't fuse (perhaps because a cmp with a memory source argument can't fuse on AMD?), and so it again points to the jne. I haven't looked at many ETW traces, so I couldn't tell you offhand - but for those who have, do the samples usually show on the expensive instructions (things like div and loads that miss are a giveaway), or on the one after? Added: Per Agner, I guess the "fusion + no skid" is the most likely (from the Ryzen section of microarchitecture.pdf): > A CMP or TEST instruction immediately followed by a conditional jump can be fused into a single μop. This applies to all versions of the CMP and TEST instructions and all conditional jumps, except if the CMP or TEST instruction has a rip-relative address or both a displace- ment and an immediate operand. That also lines up with the cmp having exactly 0 samples, unlike any other instruction of the 7: that's a common indication of fusion.