4 ms·
>The 16 or 32 architectural registers of the ISA are pretty much irrelevant compared to the capabilities of the out-of-order engine on modern chips. This can h
by _chris_ 6y ago
>The 16 or 32 architectural registers of the ISA are pretty much irrelevant compared to the capabilities of the out-of-order engine on modern chips.
This can have a big effect on memory traffic due to unnecessary moves and stack pops/pushes.
- dragontamer 6y agoAll register moves are just renames. On x86 systems like Skylake or Zen3, it doesn't even use an execution pipeline. Literally zero resources used, aside from the decoder. Heck, you can perform "xor rax, rax" as much as you'd like, because that's also a rename. It doesn't use any pipeline at all either. Register renaming (aka: "malloc" a register) is the fastest operation on modern CPUs. That's xor blah,blah, or mov foo, bar, etc. etc. ---------- Its only really an issue if you somehow need more width ("Instruction-level parallelism") than what the architectural registers provide. And even then... I'm not sure if it matters. There's store-to-load forwarding. So any register you write to L1 cache (that is read back in later on) will be store-to-load forwarded, and the whole memory read will be bypassed anyway. And even then, L1 cache is 4-clock ticks latency, and issued at the full speed of the chip's load/store units. Even if store-to-load forwarding failed for some reason, L1 cache is damn near the speed of a register.
- _chris_ 6y agoRegister moves aren't free, even if you do renaming tricks. They take up fetch bytes, they take up decode slots, they take up a lot more resources than just Regfile/ALU bandwidth.
- FullyFunctional 6y agoYou seem to misunderstand. With too few registers you may have to spill to the stack. You can't turn those into register moves in the presence of stores without a very advanced memory disambiguation.
- gpderetta 6y agotechnically Zen2 does exactly that. They do still use store/load bandwidth though. I guess that counts as very advanced memory disambiguation.
- FullyFunctional 6y agoAll high-end processors do it, but that doesn't come for free (design/verification effort + silicon area/power). Also, the capacity for this is limited to the size of the store buffer so without the needless spills, you can apply all this expensive machinery to more real memory ops. (A similarly but different debate could be had over all the save/reloads we incur on function entry/exits. 29k and SPARC's register windows were attempts at avoiding those).
- FullyFunctional 6y agoI failed to add that the stores aren't eliminated by this either so we are also incurring increased memory traffic unnecessarily.
- jcranmer 6y agoWe're talking about stores to the stack, which is likely to not be used by other threads/processors, so all of the values are being modified in a cache entry helpfully held in the Modified state and incurring no bus traffic. It will use up the traffic to/from the cache, though.
- FullyFunctional 6y agoYes I did mean cache traffic, poor choice of words, but the point is the same (filling up the store buffers etc).