4 ms·
NumPy default is that you iterate over the earlier dimensions first. The slow code is likely at least partially slow due to branch misprediction (this is speci
by itamarst 3y ago
NumPy default is that you iterate over the earlier dimensions first.
The slow code is likely at least partially slow due to branch misprediction (this is specific to my CPU, not true on CPUs with AVX-512), see https://pythonspeed.com/articles/speeding-up-numba/ https://pythonspeed.com/articles/speeding-up-numba/ where I use `perf stat` to get branch misprediction numbers on similar code.
With SIMD disabled there's also a clear difference in IPC, I believe.
The bigger picture though is that the goal of this article is not to demonstrate speeding up code, it's to ask about level of parallelism given unchanging code. Obviously all things being equal you'll do better if you can make your code faster, but code does get deployed, and when it's deployed you need to choose parallelism levels, regardless of how good the code is.
- dahart 3y ago> when it's deployed you need to choose parallelism levels, regardless of how good the code is. Yes, absolutely, exactly. That’s why it can be really helpful to pinpoint that cause of slowdown, right? It might not matter at deployment time if you have an automated shmoo that calculates the optimal thread load, but knowing the actual causes might be critical to the process, and/or really help if you don’t do it automatically. (For one, it’s possible the conditions for optimal thread load could change over the course of a run.)