3 ms·
We are observing different effects. I suggest you apply the latest microcode update, seeing that you have LSD_UOPS being dispatched; the latest update disabled
by pbsd 9y ago
We are observing different effects. I suggest you apply the latest microcode update, seeing that you have LSD_UOPS being dispatched; the latest update disabled the LSD altogether, due to the SKL150 CPU errata. The LSD was also permanently disabled in the newer Skylake-X, probably for the same reason.
In any case, I have to admit being wrong; port 6 has nothing to do with it, though the general principle of stalling the frontend to prevent the backend being overloaded still holds. For example, here's how to get the speedup without any change in control flow:
mov rcx, 400000000-1
align 32
nop
up:
align 32
add rcx, [counter]
align 32
dec rcx
align 32
jnz up
ret
The align directives here must resolve to an appropriate series of multi-byte nops. What we are doing here is to purposefully waste the uops cache. Each useful instruction comes at the beginning of a new 32-byte uop cache line, and the CPU cannot load more than 1 cache line per cycle. Effectively, we're limiting the rate of decoding by adding nops, which will not result in actual uops after decoding. I suspect the CALL and JMP instructions in the other variants were doing the same thing, but in a more obscure way. It's possible that CALL disables the LSD, which would matter in Sandy Bridge, but not in (updated) Skylake.
As to why this actually makes things faster, I don't know. There are many more stalls owing to the reservation station being full in the tight loop:
1,348,491,506 resource_stalls_any
1,694,613,980 uops_issued_stall_cycles
759,763,767 uops_executed_stall_cycles
994,974,421 uops_retired_stall_cycles
vs
109,300,935 resource_stalls_any
120,303,318 uops_issued_stall_cycles
164,322,430 uops_executed_stall_cycles
15,834,456 uops_retired_stall_cycles
But how does this translate to the observed performance penalty? One clue lies in the port breakdown:
756,686,062 uops_dispatched_port_port_0
782,579,771 uops_dispatched_port_port_1
279,387,822 uops_dispatched_port_port_2
314,214,218 uops_dispatched_port_port_3
1,994,313,938 uops_dispatched_port_port_4
807,514,909 uops_dispatched_port_port_5
1,050,884,205 uops_dispatched_port_port_6
209,316,457 uops_dispatched_port_port_7
vs
640,957,968 uops_dispatched_port_port_0
637,621,266 uops_dispatched_port_port_1
399,224,647 uops_dispatched_port_port_2
400,205,411 uops_dispatched_port_port_3
1,610,294,604 uops_dispatched_port_port_4
643,780,748 uops_dispatched_port_port_5
848,142,785 uops_dispatched_port_port_6
470,700 uops_dispatched_port_port_7
Focus on port 4, the only port available for memory stores. It does nothing else, so ideally we would expect ~400000000 uops to be dispatched to port 4. Instead, we have ~4x that for the fast loop, and ~5x that for the slow loop. These, presumably, are speculative stores that had to be invalidated. By dispatching fewer instructions that we can handle per cycle, there is less speculation and thus less wasted work?
- rayiner 9y agoThe loop stream detector can’t handle a call/ret in the loop.
- nkurz 9y agoYes. https://www.intel.com/content/www/us/en/architecture-and-technology/64-ia-32-architectures-optimization-manual.html https://www.intel.com/content/www/us/en/architecture-and-tec... Section 2.3.2.4 describes it for Sandy Bridge (and the rules don't seem to have changed in later generations): The loops with the following attributes qualify for LSD/micro-op queue replay: • Up to eight chunk fetches of 32-instruction-bytes. • Up to 28 micro-ops (~28 instructions). • All micro-ops are also resident in the Decoded ICache. • Can contain no more than eight taken branches and none of them can be a CALL or RET. • Cannot have mismatched stack operations. For example, more PUSH than POP instructions.
- nkurz 9y agoI suggest you apply the latest microcode update, seeing that you have LSD_UOPS being dispatched I haven't been following that bug closely, but I think I'm OK on this machine since hyperthreading is turned off? This a a benchmarking machine, so I have to be careful about making changes that will affect historical results. the general principle of stalling the frontend to prevent the backend being overloaded still holds I think I agree that this is what is happening. Why it's necessary, and the exact mechanism by which it helps are places that I'm uncertain about. I presume you are seeing the instructions showing up as IDQ_MITE_UOPS? From https://www.intel.com/content/www/us/en/architecture-and-technology/64-ia-32-architectures-optimization-manual.html https://www.intel.com/content/www/us/en/architecture-and-tec... Section 2.3.2.2, here's some more limitations of the Decoded ICache: The Decoded ICache consists of 32 sets. Each set contains eight Ways. Each Way can hold up to six micro-ops. The Decoded ICache can ideally hold up to 1536 micro-ops. The following are some of the rules how the Decoded ICache is filled with micro-ops: • All micro-ops in a Way represent instructions which are statically contiguous in the code and have their EIPs within the same aligned 32-byte region. • Up to three Ways may be dedicated to the same 32-byte aligned chunk, allowing a total of 18 micro-ops to be cached per 32-byte region of the original IA program. • A multi micro-op instruction cannot be split across Ways. • Up to two branches are allowed per Way. • An instruction which turns on the MSROM consumes an entire Way. • A non-conditional branch is the last micro-op in a Way. • Micro-fused micro-ops (load+op and stores) are kept as one micro-op. • A pair of macro-fused instructions is kept as one micro-op. • Instructions with 64-bit immediate require two slots to hold the immediate. As theory would predict, I was able to get the "fast" speedy by forcing legacy decoding by adding a enough junk TEST/JZ pairs. The "non-conditional branch" limitation might explain why both JMP and CALL/RET have the same effect. I wonder if the occasional alignment effects I see have to do with splitting op-stores across a 32B boundary? Focus on port 4, the only port available for memory stores. Yes, we'd expect 4B, and we instead see 1.6B for the fast case and 2.0B for the slow. My version above is at 1.4B, and I'd guess is proportionally faster. Have you seen tricks that reduce this to some lower number? We'd also expect P2 plus P3 to equal 4B? It's interesting that these are closer, but still quite a bit over. By dispatching fewer instructions that we can handle per cycle, there is less speculation and thus less wasted work? While there is less waste in terms of power efficiency, are we actually seeing P4 be a bottleneck? It's hard to tell as the total cycles also goes up, but I think it's still at 90% rather than 100% utilization. So while it looks bad, I'm not sure that anything will necessarily go faster if we reduce the number further. It may be that the maximum difference is whether we get a 4 cycle turnaround on the store-load forwarding versus a 5 cycle.