4 ms·
Oooh, I forgot about this. Only thing I can think of is something like: 1. Assume that some values can be treated internally as "physical register plus a small
by eigenform 2y ago
Oooh, I forgot about this. Only thing I can think of is something like:
1. Assume that some values can be treated internally as "physical register plus a small immediate"
2. Assume that, in some cases, a physical register is known to be zero and the value can be represented as "zero plus a small immediate" (without reference to a physical register?)
3. Since 'count' is always expected to be <64, you technically don't need to use a physical register for it (since the immediate bits can always just be carried around with the name).
4. The 3-cycle case occurs when the physical register read cannot be optimized away??
(Or maybe it's just that the whole shift operation can occur at rename in the fast case??)
- phire 2y ago> 4. The 3-cycle case occurs when the physical register read cannot be optimized away?? No... the 3-cycle case seems to be when the physical register read is optimized away. I think it's some kind of stupid implementation bug.
- eigenform 2y agoLike, the scheduler is just waiting unnecessarily for both operands, despite the fact that 'count' has already been resolved at rename?
- phire 2y agoThe 3 cycles latency casts massive suspicion on the bypass network. But I don't see how the bypass network could be bugged without causing the incorrect result. So the scheduler doesn't know how to bypass this "shift with small immediate" micro op. Or maybe the bypass network is bugged, and what we are seeing is a chicken bit set by the microcode that disables the bypass network for this one micro op that does shift.
- eigenform 2y ago> But I don't see how the bypass network could be bugged without causing the incorrect result. Maybe if they really rely on this kind of forwarding in many cases, it's not unreasonable to expect that latency can be generated by having to recover from "incorrect PRF read" (like I imagine there's also a case for recovery from "incorrect forwarding")
- phire 2y agoYeah, "incorrect PRF read" is something that might exist. I know modern CPUs will sometimes schedule uops that consume the result of load instruction, with the assumption the load will hit L1 cache. If the load actually missed L1, it's not going to find out until that uop tries to read the value coming in from L1 over the bypass network. So that uop needs to be aborted and rescheduled later. And I assume this is generic enough to catch any "incorrect forwarding", because there are other variable length instructions (like division) that would benefit from this optimistic scheduling. But my gut is to only have these checks on the bypass network, and only ever schedule PRF reads after you know the correct value has been stored.
- Bulat_Ziganshin 2y agomaybe, the bypass network doesn't include these "constant registers"? a bit like zen5 where some 1-cycle SIMD ops are executed in 2 cycles, probably for shortcomings of the same network