4 ms·
All modern high-performance processors do multiple instructions per clock.
by NohatCoder 5y ago
All modern high-performance processors do multiple instructions per clock.
- dahart 5y agoThere’s a lot of ambiguity behind that statement (fetch, decode, queue, pipeline, etc.) I’m just saying what I read in the article, the peak bandwidth calculation was making the assumption of a peak throughput of one instruction per cycle, according to the author. The best answer might be to profile it yourself and see what happens on your processor, it’s only a few lines of code and looks incredibly easy to do.
- NohatCoder 5y agoThere is usually a bunch of weird specifics when it comes to calculating performance, but this one happens to be quite simple as we bottleneck on 1 write per clock, and the add instruction latency of 1, so none of the hairy stuff should come into play. Not knowing what cpu compiler and settings are used I can't do much to replicate the test.
- dahart 5y agoIt would be compelling enough to use the author’s 3-line loop as-is on your processor with your own C or C++ compiler. Or the compiled baseline assembly from the article…
- sereja 5y agoIt's called superscalar processing. CPUs can execute more than one instruction concurrently on each pipeline stage if there is no dependency between them. In the prefix sum, we can fetch and add the next element simultaneously with the accumulator of the previous iteration is being written back. https://en.algorithmica.org/hpc/pipelining/ https://en.algorithmica.org/hpc/pipelining/
- dahart 5y agoThat link shows 1 fetch per cycle. “Pipelining does not reduce actual latency”. So your comment fails to explain both multiple fetch per cycle and higher than 1 instruction per cycle throughput on average, which is what we were discussing above.