4 ms·
I was with you up until the last paragraph. Taken literally you seem to suggest there is no such thing as a CPU-bound workload. That's obviously not the case (c
by _dps 9y ago
I was with you up until the last paragraph. Taken literally you seem to suggest there is no such thing as a CPU-bound workload. That's obviously not the case (cryptography is just one such example), but I would agree that many people think they are CPU-bound when they are really constrained by something else.
Secondly, Python and the patterns its expressiveness encourages are terrible for cache performance. In a simple C program it's easy to do something non-trivial in the space provided by L1 cache — in Python it's quite difficult even to reason about what's going to be in L1 if you're using any of the fancy features.
- smitherfield 9y ago>in Python it's quite difficult even to reason about what's going to be in L1 The interpreter's stack?
- _dps 9y agoI haven't looked at it in a while, so I could be wrong, but I think with small enough programs you can still squeeze some payload into L1 in long tight loops where you're not jumping up and down the Python stack a lot. But your overall point stands: if you're writing non-trivial Python programs your L1 is usually spent on language/runtime overhead.
- kevin_thibedeau 9y agoWhen such things matter you drop to Cython and avoid interacting with PyObjects. Then you get native performance for tight loops.
- theli0nheart 9y agoGood point. I see what you're saying and I was definitely not suggesting that. I was speaking mainly from a web application perspective, where program execution on the server is very low on the list of items affecting application speed as perceived by the end-user. Scientific computing, on the other hand, is a completely different animal.