4 ms·
(note: this comes mostly from a background in scientific/numerical programming in C) You are certainly right if one takes "predictable" to mean something close
by _dps 12y ago
(note: this comes mostly from a background in scientific/numerical programming in C)
You are certainly right if one takes "predictable" to mean something close to "deterministic"; on the other hand I think there are numerous almost-deterministic characteristics that can be evaluated without resorting to plain-ol' empiricism (particularly involving likelihood of hits in L1/L2/L3 and likelihood of being able to speculatively execute particular branches of code).
So I would agree that my a priori predictions may be a factor of 2-3x off; OTOH I often succeed in predicting that a particular tight loop will never have to leave L1, or that a low-entropy condition in the loop will be made essentially irrelevant by speculative execution.
With that proviso, I share your belief that JITs seem to be the best hope for top performance (and I am a heavy user of LuaJIT).
- seanmcdirmid 12y agoIf you use CUDA + a GPU, you know for sure, since you are basically responsible for scheduling all memory movements yourself. On the other hand, even CUDA is doing some optimizations behind the scenes that can drastically effect the performance of your code in ways that are not so obvious. A lot of the HPC work has moved over to GPUs, it is amazing what one can do when you have almost complete control over the memory hierarchy.
- pron 12y agoMemory and caching is a big part of nondeterminism, but it's not just about memory. GPUs employ a far more primitive execution strategy than modern server CPUs with their branch prediction, ILP etc. Also, GPUs, while terrific for parallel workloads, are terrible for concurrent workloads, which require very intricate branching.
- seanmcdirmid 12y agoWell, you can branch as much as you want on the GPU, as long as your branches are coherent among each thread sharing the same control unit :) I wouldn't say this is "primitive", just a very different way in thinking about computation that has been very very very effective for a lot of use cases. GPUs are obviously not the solution for concurrent or irregular workloads (yet), but many are surprised how much mileage one can get out of Python + CUDA for scientific workloads. The only point I was making is that there is a world where the hardware is much more deterministic (even if most of us can't go there).
- pron 12y agoNo argument there. BTW, for anyone interested in an overview of the non-determinism built underlying modern hardware architectures, I recommend watching this great talk[1] -- A Crash Course in Modern Hardware -- by Cliff Click, one of the world's top JIT experts. [1]: http://www.infoq.com/presentations/click-crash-course-modern-hardware http://www.infoq.com/presentations/click-crash-course-modern...