3 ms·
1.3-1.4x slowdown is a lot more than I expected (I know it's for synthetic benchmarks but still...) Can someone explain (or link to an article) how a tweak to
by _cs2017_ 8y ago
1.3-1.4x slowdown is a lot more than I expected (I know it's for synthetic benchmarks but still...)
Can someone explain (or link to an article) how a tweak to HT branch prediction heuristic can have such a huge impact on performance?
- janoc 8y agoIt is unlikely the branch prediction heuristics that is the problem (that in the Intel's microcode). The problem is in the mitigations necessary to make it impossible/more difficult to exploit these side channel attacks. And that is costly because memory and needs to be moved around constantly. So that adds a ton of extra overhead whenever a context switch is made.
- BeeOnRope 8y agoThe impact is big enough that one would suspect the microcode simply disables indirect branch prediction, so you pay a 16-20 cycle penalty per branch. Indirect branches just aren't frequent enough to explain such a regression via say a simple reduction in prediction resources. I can test it once I get the new firmware.
- jcelerier 8y ago> Indirect branches just aren't frequent enough aren't vtable calls / function pointer calls indirect branches ?
- BeeOnRope 8y agoYes, they are (at least when the compiler cannot devirtualize them) - but they make up a fairly small fraction of the total instructions in a typical program - and probably very small in something like cinebench, which also showed a big regression.
- jcelerier 8y ago> but they make up a fairly small fraction of the total instructions in a typical but if their cost increases by a large factor... besides, in any large compiled program, the core would certainly be based around some kind of programmable pipeline, and these would generally be implemented like this unless they wrote their own JIT compiler.
- BeeOnRope 8y agoYes - but for their cost to increase by such a large factor, the only obvious thing I can think of is that their prediction is disabled. I didn't follow your comment about a "programmable pipeline". I don't think many or any of the Phoronix benchmarks are based on a pipeline with indirect branches at their core.
- jcelerier 8y ago> I don't think many or any of the Phoronix benchmarks are based on a pipeline with indirect branches at their core. I think a bunch are. e.g. for instance FFMPEG / libavfilter which is basically a node graph set up at runtime. Don't know for cinebench since it's closed source, but Blender present in the benchmarks is also based around a nodal rendering architecture. Stuff like PHP / CGI also heavily depend on function pointers for their behaviours - PHP with its plugin architecture, and CGI where all web requests go through FPs : https://github.com/php/php-src/blob/master/main/fastcgi.c#L880 https://github.com/php/php-src/blob/master/main/fastcgi.c#L8....
- BeeOnRope 8y agoRight, I think I understand what you are saying about "pipelined" implementations. Sure, I can believe that at a high level there are some indirect branches to implement some kind of processing pipeline: but you'd have thousands or millions of instructions doing the heavy lifting for each chunk of data that passes though the pipeline, for every branch that you need to take to get to the next stage. So I still doubt that that indirect branches are "dense" in those benchmarks: it just doesn't make sense since the core work they are doing are highly tuned encode/decode/whatever kernels, even if there is a control layer over top of that using indirect branches.