5 ms·
My understanding is that they do not predict the target of the next branch but of the one after the next (2-ahead). This is probably much harder than next-branc
by cpldcpu 2y ago
My understanding is that they do not predict the target of the next branch but of the one after the next (2-ahead). This is probably much harder than next-branch prediction but does allows to initiate code fetch much earlier to feed even deeper pipelines.
- emn13 2y agoAh, that makes sense in the context of the article - thanks!
- layer8 2y agoSurely you must also predict the next branch to predict the one after. Otherwise you wouldn’t know which is the one after. Given that, I still don’t understand how predicting the next two branches is different from predicting the next branch and then the next after that, i.e. two times the same thing.
- ithkuil 2y agoInterestingly you can build a branch predictor that predicts the second branch without predicting the first. A branch predictor result is just a tuple of ("branch instruction address", "branch target address") that hints the processor that when the CPU will encounter a given branch instructions in the future (at "branch instruction address") it will likely branch to the branch target and so it would make sense to start fetching that address and filling the instruction pipeline with whatever steps are safe to perform before the jump will be actually performed. Now, commonly this branch happens to be at the end of the current basic block and I assume some branch predictors may also leverage this fact in order to encode only offsets from the current instruction pointer. But there is no reason why the branch location might be after some other branches may be taken. As long as the cpu eventually gets to that branch location the prediction will be useful. If the IP never reaches that location it's like the branch was never actually taken.
- Filligree 2y agoBuilding on the sibling comment: if (a) { ... } if (b) { return x; } else { return y; } The two branches can be wholly independent, but predicting the second is still a two-ahed prediction.
- toast0 2y ago> Given that, I still don’t understand how predicting the next two branches is different from predicting the next branch and then the next after that, i.e. two times the same thing. I'm not involved in CPU design, I just read a lot, but... I think you need to do something special to have a second prediction, because you have to track three windows of out of order execution: Window 0: code you're definitely running but is still being completed. Window 1: code from the branch you think will be taken Window 2: code from the 2-ahead branch you think will be taken. If you figure out that the window 1 branch isn't taken, you have to drop the whole pipeline (pipeline bubble). But if you figure out that window 1 is taken, then window 1 becomes window 0 and window 2 bcomes window 1. With a 1 ahead predictor, the pipeline stalls if you get to a conditional branch while speculating in window 1, because the processor can't manage three instruction windows. IMHO, it sounds like if the core is doing SMT and both threads are active, each thread only gets 1-ahead prediction because the two windows are statically divided between the cpu threads. This may mean a) a significant boost for some loads when SMT is not in use and b) SMT can branch predict in both threads in the same cycle, I don't think that was possible on AMD before (no idea for other vendors)
- gpderetta 2y agoWith a typical OoO CPU you have reorder buffers in the hundreds of instructions, so you are not speculating two or three branches but potentially dozens ahead of non-speculative execution (and that's why prediction accuracy is so important). So the question remain, what's the innovation here? I'm sure there is something, but it is not simply speculating two ahead. I need tor was further, bit this might be an optimization in the fetched tage to avoid a bubble when fetching consecutive taken branches.
- fulafel 2y ago> Surely you must also predict the next branch to predict the one after. Otherwise you wouldn’t know which is the one after. I'd think if you are at PC N and there are branches at N+1 and N+2, predicting just branch N+2 is fine because you predicted the N+1 branch previously, at PC N-1.
- flamedoge 2y agoI wonder what they need before this change. Branch predictor hardware may not have accounted for depth beyond single conditional branch? but pipeline was probably always filled, unpredicted.