4 ms·
I'll have to repeat my question in other words. How many MIPS do you think, remain, if every third instruction is a mispredicted branch and every second memory
by static_noise 10y ago
I'll have to repeat my question in other words. How many MIPS do you think, remain, if every third instruction is a mispredicted branch and every second memory access is a cache miss?
Modern CPUs use pipelining which executes many instructions parallel. This only works well if everything goes as predicted. If you have an algorithm which works contrary to what the branch prediction thinks and a cache which does not hold the data you need, your performance goes down the drain. Those MIPS mean nothing if not put into the right context.
- Dylan16807 10y agoYou can do a lot to work around branches. Cache misses make CPU speed irrelevant, but when you look at your memory system it's another instance of the same problem with the same tradeoffs of frequency vs. work per cycle. And when the "right context" is waiting for the outside world, that's not the most important context.
- twoodfin 10y agoOne of my favorite papers is relevant here: http://www.hpl.hp.com/techreports/Compaq-DEC/WRL-93-6.pdf http://www.hpl.hp.com/techreports/Compaq-DEC/WRL-93-6.pdf It explores the limits of instruction-level parallelism. If you had a processor that could dispatch an unlimited number of independent instructions simultaneously, how much of an improvement in common algorithms would you get?
- static_noise 10y agoThey get an average parallelism of 4-10 with some standard algorithms assuming unlimited resources. Where are we now with modern intel processors on those algorithms? 2-3?