3 ms·
The problem is that unlike data prefetch, the so-called "instruction prefetch" is not actually prefetch at all. It's simply speculative execution. Which doesn'
by phire 25d ago
The problem is that unlike data prefetch, the so-called "instruction prefetch" is not actually prefetch at all.
It's simply speculative execution. Which doesn't look any different to regular execution. The fetcher has no idea that its predicted branch is about to invalidated and flushed, otherwise it would never have issued that fetch.
Actually, on a modern OoO core, [0] it's very rare for the instruction fetcher to not be doing speculative fetches. Even when it's not predicting a branch, the fact that it has "predicted" the lack of a branch is speculative in itself. It assumes it didn't fetch a branch in the last cycle, but it can't be sure until after instruction decoding, which takes at least 2 cycles (more on larger L1i caches).
About the only time the instruction fetcher is not doing speculative fetching is for a single cycle after each miss-predicted branch.
[0] Or even something technically in-order, like the Cortex A53 cores here. They might issue in-order, but because of how they implement dual issue, they look somewhat close to a simple OoO core... I suspect they actually do register renaming. And (most importantly) importantly they have a branch predictor.
- repiret 25d agoNo, the problem really is the prefetch itself when there are undesirable side-effects if the bus sees that memory read.
- phire 25d agoWell yes. That is why the "prefetch" is a problem. But the original question was asking why disabling data prefetching to a memory region didn't automatically disable instruction prefetching at the same time. And the answer is that speculative execution is a completely different mechanism that I'm not even sure can be disabled, at least not per memory region.
- renox 25d agoWhy do you say this? The post is quite clear that marking the memory region as NX fix the issue caused by speculative execution.
- sleirsgoevy 25d agoLooks like you know some internal details of the A53 cores, can you share the source?
- phire 25d agoThe fact it does speculative execution and dual-issue is well documented. The chipsandcheese article [0] is probably the best overview. The "how" it does dual-issue is not documented at all, so I'm speculating. There is no smoking gun saying "register renaming". While it would be possible to implement its known capabilities without any kind of renaming, it would be so much simpler to implement it with register renaming. The thing is.. once your forwarding network and hazard detection gets complicated enough, it basically becomes a janky form of register renaming. So it's cleaner to just implement proper register renaming, and actually saves hardware. While the A53 might issue in-order, the pipelines have different lengths and they don't finish executing in order. I find the fact the A53 can speculate a few instructions past a cache miss to be very interesting, along with the fact it can issue two writes to the same logical register in a single cycle. Also, the smoking gun is that the A510 (same lineage as the A53) is documented to do out-of-order issue (see chipsandcheese again [1]), so it must be doing register renaming. ARM still insist on still calling it an in-order core because it's OoO is so much more limited to modern OoO cores, but it's more OoO than early PowerPC designs (including the G3) that everyone is happy applying the OoO label to. [0] https://chipsandcheese.com/p/arms-cortex-a53-tiny-but-important https://chipsandcheese.com/p/arms-cortex-a53-tiny-but-import... [1] https://chipsandcheese.com/p/arms-cortex-a510-two-kids-in-a-trench-coat https://chipsandcheese.com/p/arms-cortex-a510-two-kids-in-a-...