12 ms·
> each sum supposedly need 0,1ns Sub-cycle latency for an addition would be quite amazing... Nevertheless your point remains. > 100ns=0,1 ms 0.1μs, not 0.1ms
by _hl_ 6y ago
> each sum supposedly need 0,1ns
Sub-cycle latency for an addition would be quite amazing... Nevertheless your point remains.
> 100ns=0,1 ms
0.1μs, not 0.1ms
- andi999 6y agoOf course! Even worse. I need my coffee (and stop yelling at the students when they mess this up)
- titzer 6y agoModern superscalar processors have more than one execution unit. The latest Intel microarchitecture has 10 execution ports, which means it can theoretically execute 10 μ-ops per cycle. With a 200+ entry reorder buffer, the processor is trying hard to cram those execution units full every cycle. How successful it is depends on dynamic data dependencies and fetch/issue bandwidth. At 3ghz, to reach 0.1ns latency (on average), it only needs to achieve 3 instructions per clock.
- andi999 6y agoI fully agree with your post. I am just not sure if you are trying to make a point for larger datasets in benchmark (like I did) or you are trying to imply that this is not necessary?
- titzer 6y agoI generally agree the benchmark is too small, but I have serious issues with Lemire's approach to benchmarking in general. Things get pretty nuts at that small of a timescale and things like frequency scaling can make a huge difference. I would draw absolutely zero conclusions on these reported results.
- ggrrhh_ta 6y agoI had that sort of impression as well, but after playing with the benchmark and seeing the assembly code, and checking the actual theoretical throughputs for my microprocessor, I realized that in this case, and previous ones, there was nothing relevant to the results that had to do with the way benchmarking was implemented. These pieces of code would be thought to be in the innermost loop, and practically they might not actually process a very large part of the array. There are solutions, and better approaches, sure, but they pinpoint real issues that cause spread in your performance across microprocessors & compilers that are worth having in mind.
- _hl_ 6y agoRight, but afaik only one instruction per 'hyperthread' is fetched every cycle, but I'm happy to be corrected here.
- titzer 6y ago5-6 per cycle per thread is more typical. Can be faster with a micro-op cache or a loop stream buffer.
- Dylan16807 6y agohttps://en.wikichip.org/w/images/7/7e/skylake_block_diagram.svg https://en.wikichip.org/w/images/7/7e/skylake_block_diagram.... Way more than that.
- andi999 6y agoI fully agree these things are difficult and not well defined how to benchmark this. Actually thinking on what you wrote the sum+= should lead to WAW and RAW conflicts, which might not allow full ultilization of the execution ports. It would be interesting to see if having 10 temp(or 16ä ) variables each accumulating one tenth of the dara will run faster. All conflicts should be solvable then