4 ms·
Author here. > I can see why they described it as a '5 entry load buffer' for instance but it's not really accurate to the true microarchitecture. That's what
by clamchowder 3y ago
Author here.
> I can see why they described it as a '5 entry load buffer' for instance but it's not really accurate to the true microarchitecture.
That's what it looks like to software. I can put 5 loads between two other loads that miss cache, and the cache miss latency will overlap. It feels like a 5 entry load buffer to software. I'm more than happy to describe the true microarchitecture if Arm talks about it :) Otherwise it'll be "hey, how does it look to software from a performance perspective?"
> There's also lots of fun details of the crazy things you have to do to hit your frequency target (often where you just ignore certain rare hazard or error conditions and detect them later and fix it up).
Yeah people do this everywhere. Sometimes you can find 20+ cycle penalties for stuff like a load that depends on a page-crossing store. The trick is making sure the hazards are indeed rare relative to the penalty. Also the biggest penalties in practice are cache misses and DRAM latency in the hundreds of cycles. Intra-core penalties never come close.