5 ms·
It takes 1 read port for 1 clock. Any AVX2-capable processor has at least 1 read and 1 separate write port, therefore we should get 1 processing cycle per clock
by NohatCoder 5y ago
It takes 1 read port for 1 clock. Any AVX2-capable processor has at least 1 read and 1 separate write port, therefore we should get 1 processing cycle per clock, the reported number is only around a third of that.
- deleted 5y ago[deleted]
- dahart 5y agoThe author assumes an average of (slightly over) 2 instructions per iteration, which means the theoretical limit is a half iteration per clock, assuming no cache misses ever. How can you reduce that to 1 processing cycle per clock in a single thread, is there a load-add-store instruction, or a way to get multiple instructions per clock through a single thread?
- NohatCoder 5y agoAll modern high-performance processors do multiple instructions per clock.
- dahart 5y agoThere’s a lot of ambiguity behind that statement (fetch, decode, queue, pipeline, etc.) I’m just saying what I read in the article, the peak bandwidth calculation was making the assumption of a peak throughput of one instruction per cycle, according to the author. The best answer might be to profile it yourself and see what happens on your processor, it’s only a few lines of code and looks incredibly easy to do.
- NohatCoder 5y agoThere is usually a bunch of weird specifics when it comes to calculating performance, but this one happens to be quite simple as we bottleneck on 1 write per clock, and the add instruction latency of 1, so none of the hairy stuff should come into play. Not knowing what cpu compiler and settings are used I can't do much to replicate the test.
- dahart 5y agoIt would be compelling enough to use the author’s 3-line loop as-is on your processor with your own C or C++ compiler. Or the compiled baseline assembly from the article…
- sereja 5y agoIt's called superscalar processing. CPUs can execute more than one instruction concurrently on each pipeline stage if there is no dependency between them. In the prefix sum, we can fetch and add the next element simultaneously with the accumulator of the previous iteration is being written back. https://en.algorithmica.org/hpc/pipelining/ https://en.algorithmica.org/hpc/pipelining/
- dahart 5y agoThat link shows 1 fetch per cycle. “Pipelining does not reduce actual latency”. So your comment fails to explain both multiple fetch per cycle and higher than 1 instruction per cycle throughput on average, which is what we were discussing above.