4 ms·
I'm rather surprised at the claim that "but it might easily execute 7 billion instructions per second on a single core". I'd even question it except the author'
by tempguy9999 7y ago
I'm rather surprised at the claim that "but it might easily execute 7 billion instructions per second on a single core". I'd even question it except the author's an expert.
If you can keep it fed then ok but one cache miss to main mem, either instruction or data, will allow the instruction buffers to completely empty and stay empty for quite a long time. I don't think you can control placement to reasonably assure cache hits always for anything but the most trivial code, am I missing something?
Also if you could keep a consistent throughput like this I wonder if thermal throttling might have to kick in. I mean you're doing a lot of work...
- touisteur 7y agoI can't find it back but in a recent article I read that it was useful to have an idea of the upper-boundary abilities of an arch+algorithm, so that you 'know' what you're aiming for, but it might not be attainable practically without huge human or decades of superoptimizer effort... Yes if your algorithm reaches for cold data, you'll get hit. Can you get around that? Do you really need to hit the cache when you're computing the seven-billionth decimal of pi or factoring numbers ? This work is quite interesting, if only for compilers or superoptimizers.