3 ms·
Really great article and comments :) Just as a point of reference, I currently see around 1.25 - 1.5 op/cycle on carefully crafted highly parallel lock-free, s
by socmag 9y ago
Really great article and comments :)
Just as a point of reference, I currently see around 1.25 - 1.5 op/cycle on carefully crafted highly parallel lock-free, stall-free code... say running on 8 threads. Code that has 0.01% branch misprediction.
Unfortunately in my case as others mention, access to the I/O ports and memory latency, is a real limiting factor. The CPU is just... waiting
Getting to the Holy Grail that Ryg talks about of 3 instructions per cycle is really hard with non-vectorizable workloads - like screwing around with hash tables that have no chance of fitting in L1/L3, and not being able to really make much use of SIMD, even if you are paying attention to cache-lines.
Most apps barely scrape by at 0.5 instructions/cycle or worse and spend most of the time bouncing on the kernel for stupid stuff. Not good.
Absolutely <3 performance freaks!