4 ms·
In this CPU, AMD first introduced the "stack engine". This is likely a bug in that feature. It has a speculative stack address delta register in the front-end t
by throwawaylinux 4y ago
In this CPU, AMD first introduced the "stack engine". This is likely a bug in that feature. It has a speculative stack address delta register in the front-end that is updated directly with push/pop instructions, and that delta is dispatched with the stack memory uop to be added to the original stack address register when doing address generation in the load/store units.
The delta has to be small, because you don't want big adders in the front end and because the delta has to be sent down with the push/pop memory uops. That means it can overflow or underflow, at which point it has to be reset by sending a synchronize operation to the back-end to update the original stack register (agner has a better description).
So the delta register is probably 10-12 bits on Barcelona, and this bug is probably a corner case where the stack register update is happening, hence 1024 bytes off. Perhaps there is a window where a uop can get the old delta value, but the new base value (or vice versa) when a sync is operation is concurrent.
Setting that MSR value possibly disables that stack engine feature entirely, or it could be it disables some aggressive and complicated detail where the bug is, e.g., allowing stack operations to run concurrently while flush operations are in progress.
It's not a coincidence there just happens to exist a way to disable this at runtime. The way processors are designed means that everything must be able to be observed, debugged, and fixed in the field. That means everything has to have fine-grained ability to control, disable, enable safer fallback paths, and even engage additional logic to reduce the state space in some cases (e.g., serialize pipeline while a particular operation occurs).
Usually these bug fixes decrease performance (except in cases where a performance bug is found and the fix actually increases performance), so you want the switches to be very fine-grained. So it's possible they fixed the stack engine bug without disabling it entirely.
It would be like shipping software and providing support and bug fixes for it for the next 5-10 years without patching the software, only updating the config file. It's quite amazing. For every one of these issues that hits the field, there will be many found internally during the internal hardware bring up and verification (which will be ongoing for at least part of the life of the CPU).
- rcxdude 4y agoYup, these are often colloquially called "chicken bits" because you put them in if you're afraid your new feature won't work (It will generally just fall back to the previous battle-tested implementation). I often wonder how far back you can pare a modern CPU with these.