3 ms·
Care to explain? An iterator is a nice high level concept, but the CPU still has to do the &in + i + offset arithmetic. I don't see how replacing `i` with synta
by dahfizz 3y ago
Care to explain? An iterator is a nice high level concept, but the CPU still has to do the &in + i + offset arithmetic. I don't see how replacing `i` with syntactic sugar changes the need to check for overflow.
- nostrademons 3y agoI think the point the GP is making is that with an iterator protocol, the iterator implementation itself is free to make a different choice on implementation strategies, based on the shape of the data and the hardware available, and this is transparent to client code. So for example, a container containing only primitive ints or floats on a machine with a NVidia Hopper GPU might choose to allocate arrays as a multiple of 64, and then iterate by groups of 64, taking advantage of the full warp without needing any overflow checks. Obviously a linked list or an array of strings couldn't do this, but then, they wouldn't want to, and hiding the loop behind an iterator lets the container choose the appropriate loop structure for the format and hardware. I've heard criticisms of C and C++ that they are simultaneously too high-level and too low-level. Too high-level in that the execution model doesn't actually match the sort of massively parallel numeric computations that modern hardware gives, and too low-level that the source code input into the compiler doesn't give enough information about the real structure of the program to make decisions that really matter, like algorithm choice. It's interesting that the most compute-intensive machine learning models are actually implemented in Python, which doesn't even pretend to be low-level. The reason is because the actual computation is done in GPU/TPU-specific assembly, so Python just holds the high-level intent of the model and the optimization occurs on the primitives that the processor actually uses.
- PaulDavisThe1st 3y ago"It's interesting that a lot of performance-critical code tends to be written in C++, which sometimes pretends not to be that low level. The reason is because the actual performance critical code is really running CPU-specific assembly, so C++ just holds the high level intent of the model and the optimization happens on the primitives that the processor actually uses."
- dahfizz 3y ago> the iterator implementation itself is free to make a different choice on implementation strategies That's just UB with more steps. What will the spec say? "Behavior of integer overflow is undefined. Unless the overflow happens within an iterated for loop, in which case the behavior is undefined and the iterator can do whatever it wants". > I've heard criticisms of C and C++ that they are simultaneously too high-level and too low-level. I've heard this as well, and I think there is some truth to it, but C is the least-bad offender relative to any other language. C maps extremely well to assembly. The fact that assembly no longer perfectly captures the implementation of the CPU has nothing to do with C. Every other general purpose[1] language has to target the same abstraction that C does. Given that reality, C in fact maps better to the hardware than any other language. Because it is faster than any other language. Any higher level language that gives the compiler more information about algorithm choice is slower than C is. That's the bottom line. [1] This is ignoring proprietary, hardware specific tools like CUDA. That's clearly in a different category when discussing programming languages, IMO.
- nostrademons 3y agoThe reasoning behind the decision to make integer overflow UB changes. As the thread starter mentioned, that reasoning was loops, so you don't need an overflow counter for everyday loops. Take loops out of the equation, and take certain high-performance integer computations where arguably you should have a dedicated FixedInt type, and the logical spec behavior might be silent promotion to BigInt (like JS, Python 3, Lisp, Scheme, Haskell) or a panic (like in Rust). > [1] This is ignoring proprietary, hardware specific tools like CUDA. That's clearly in a different category when discussing programming languages, IMO. Arguably they should be part of the conversation. One main reason for the recent ascendancy of NVidia over Intel is that they're basically unwrapping all the layers of microcode translation that Intel uses to make a modern superscalar processor act like an 8086, and saying "Here, we're going to devote that silicon to giving you more compute cores, you figure out how to use them effectively."
- commonlisp94 3y ago> implemented in Python, A program which constructs an AST out of python classes and spits out GPU code is a compiler. The python is never executed. Trivially, compilers can generate code faster than their hose language, but that doesn't make the host language fast. The compiler would be even faster if it were written in C++.
- nostrademons 3y agoThat's sort of the point. If you want to make programs really fast and really expressive, your target language ought to be as close to the hardware as possible, and your source language ought to be as expressive as possible, and then you just write a compiler to translate between them.