4 ms·
> I presume you are seeing the instructions showing up as IDQ_MITE_UOPS? No, the legacy decoder is virtually unused. This loop was designed to take advantage o
by pbsd 9y ago
> I presume you are seeing the instructions showing up as IDQ_MITE_UOPS?
No, the legacy decoder is virtually unused. This loop was designed to take advantage of the rules of the DSB. If you have the LSD enabled I don't quite know what the effect is.
EDIT: Sorry, just realized I screwed up translating AT&T to Intel---add rcx, [counter] should clearly be add [counter], rcx.
> It may be that the maximum difference is whether we get a 4 cycle turnaround on the store-load forwarding versus a 5 cycle.
This is my guess as well, though I have nothing to show for it. It's actually possible to bring it down to 400000000 uops by delaying the loop even further with the same technique, but it won't speed anything up. The 4-cycles-per-iteration barrier seems insurmountable, which makes sense.
- nkurz 9y agoNo, the legacy decoder is virtually unused. This loop was designed to take advantage of the rules of the DSB. OK, think I see now. The DSB path is only able to inject one "way" per cycle, and through NOP padding you are able to ensure that each "way" only contains a single instruction? While I can find clear descriptions of what can fit in a "way", I haven't found a clear statement of the "one way per cycle" restriction, although it would explain observed behavior. Practically, it seems like the effect would be the same if one "defeated" the DSB and came up with a similar strategy for causing the legacy MITE to trickle in instructions. The overall goal is to restrict speculative execution of loads and stores in cases where one knows that the speculation is going to fail. You'd think that there would be an MFENCE/LFENCE/SFENCE way of solving this more directly, but I haven't found it yet. Or maybe just a LOCK? Maybe a creative false dependency to keep the processor from getting to far ahead? Interesting. Thanks for helping explore this.
- pbsd 9y agoI later realized that padding to 32 bytes didn't make much of a difference, it is the number of vacuous NOP opcodes generated that matters. So it can be simplified to mov rcx, 400000000-1 up: nop nop nop add [counter], rcx nop nop nop dec rcx nop nop nop jnz up ret which probably works fine for the majority of decoding methods, be it MITE, LSD, or DSB.