9 ms·
That's not it. I updated the article with an experiment of processing 5 items at a time. The fast function doing 5 images at a time is slower than the slow func
by itamarst 3y ago
That's not it. I updated the article with an experiment of processing 5 items at a time. The fast function doing 5 images at a time is slower than the slow function doing 1 image at a time (24*5 > 90).
If your theory was correct, we would expect the optimal number of threads for the fast function processing 5 images at a time to be similar to that of the slow function processing 1 image at a time.
In fact, the optimal threads in this case (5 images at a time) was 20 for slow function, 10 for fast function, so essentially the same as the original setup.
- rewmie 3y ago[dead]
- bogwog 3y agoIs it a caching thing? The slow version seems less cache efficient, so if it is waiting due to cache misses, that could create an opportunity for something else to get scheduled in.
- gmm1990 3y agoI doubt it the slow version uses division instead of bit shifting. My guess would be the fast version saturated like i/o or some non cpu portion of the processor and the division one was bottle necked by the division logic in the processor.
- bogwog 3y agoBut it's iterating through the result vectors twice, so that's basically guaranteed to miss. Moving the threshold check into the loop above would at least eliminate that factor. Maybe division vs bit shifting does play a factor, but it's hard to compare that while the cache behavior is so different.
- cogman10 3y agoRecalibrate how you feel about division and multiplication. It turns out, integer division on new processors is a 1 cycle process (and has been for a while now). Most of the multicycle instructions now-a-days are things like SIMD and encryption.
- jandrewrogers 3y agoWhich processor has 1-cycle latency for integer division? Even Apple Silicon, which has the mostly highly optimized implementation I am aware of appears to have 2-cycle latency. Recent x86 are much worse, though greatly improved. Integer division is much faster than it used to be but not single cycle. Also, most of those SIMD and encryption instructions can retire one per cycle on modern cores, but that isn't the same as latency.
- cogman10 3y agoYou are correct, I mistakenly thought it was 1/cycle because I had previously remembered IMUL taking 30 cycles (and now it has a 1 cycle throughput). Agner Fog is reporting anywhere from 4->30 cycles on semi-recent architectures.
- gpderetta 3y ago2-cycle latency for division seems extremely unlikely. From the instruction tables it seems that firestorm has 7-9 cycles latency SDIV (which is excellent). It also has an impressive 2-cycles reciprocal throughput.
- jandrewrogers 3y agoQuite possible. I am not confident Apple Silicon has 2-cycle div latency, that seems improbably fast to me, but I had heard some reasonably well-sourced rumors to that effect. I have not measured it myself. Even at somewhat higher latency it is still fast enough to not be worth optimizing around in most cases, which is great.
- kukkamario 3y agoBased on the graph the fast function runtime is really short. You might be just seeing effects of efficiency vs performance cores. Lower thread count makes most of the stuff run on performance cores and task end times align more nicely. With larger number of threads tasks running on performance cores complete first and you are left waiting for efficiency tasks running on efficiency cores to complete, or context switches have to be done and tasks are moved between cores which causes overhead. You could try what happens if you have 10 times more images when running fast function. Also you have just 8 physical performance cores and 4 physical efficiency cores. Performance cores have hyper threading so they act as 2 logical cores but that doesn't mean that they can actually execute 2 threads at maximum performance. If processing tasks use the same parts of the processor core, then processor cannot run both threads at the same time and IPC will suffer. Slow task maybe uses more varied parts of the core which allows better IPC with hyper threading. So that also may reduce optimal thread count.