3 ms·
> There are many things that can be inferred from the code without needing to execute it. The question is how difficult it is to make such an inference: in one
by pron 19d ago
> There are many things that can be inferred from the code without needing to execute it. The question is how difficult it is to make such an inference: in one scenario, the compiler might attempt to track whether specific data changes-and, if it can prove this, mark the data as immutable and apply certain optimizations-whereas in another, it might already possess the information that the data is immutable.
Yes, and the important point is that when it comes to knowing things statically, abstraction and optimisation are in conflict. The whole point of abstraction is that the implementation details aren't known. So in C++ we always suffer from this problem called "zero overhead abstractions" or "abstraction costs", which means that to give the compiler the information it needs, we have to use less general abstractions, which are viral and harm evolution. What a JIT does is allow the compiler to learn the very things that abstraction hides; yes, it's a virtual call, yes, it could target anything, but I've seen it hit the same target 1000 out of the last 1000 times, so I speculate that this will continue and I'll inline even though I could be wrong.
> The same applies to the GC: the compiler can perform more optimizations when it knows when memory needs to be cleared
I understand why this could be true in theory, but in practice the problem is:
1. not that the compiler knows when an object is unreachable, but that the generated code has to do something at that point, and
2. the most efficient known memory management algorithms - moving collectors and arenas, both work in nearly the same way - are entirely predicated on freeing memory in bulk and on not doing anything when an object becomes unreachable, and so the knowledge of when an object becomes unreachable doesn't help them.
So it is true that C and C++ and Rust always statically know when an object is dead, and you could say that hypothetically they don't need to do anything with that information, but in practice they all act on that information immediately and that's inefficient.
> There remain a small number of cases, such as `switch` statements - where one branch executes 99% of the time, while the other 99 branches execute only 1% of the time.
So the main practical benefit of a JIT isn't that at all, but that it can do the "mother of all optimisations" - inlining - far more aggressively. Inlining is important because it cracks open the abstraction boundary of the inlined subroutine, and allows the compiler to further specialise and optimise things, now with the appropriate context.
Anyway, all of these fundamental questions and differences between languages with more statically known information and figuring out "unprovable" information in practice were very well known before the JVM was built to address the performance problems we had suffered from in large C++ programs. So we can argue over which workloads are helped by this and which aren't, but there is no way to say which is usually faster in the absract (because, again, these considerations were known and taken into account). It's merely an empirical question, and not one that's easy to settle. After more than 25 years of working with C++ and almost 20 years of working with Java, my default is that low-level wins on performance (if written by experts) in smaller programs, and Java wins on performance in larger programs, but of course, there are many caveats in either direction.
- someone_19 18d ago> The whole point of abstraction is that the implementation details aren't known. I disagree with that phrasing; it is better to say that abstractions allow a programmer to ignore unimportant details. For example, when developing two modules (possibly even by different teams), all they know about each other is a lean interface, without any implementation details. However, the compiler might know everything. > So in C++ we always suffer from this problem called "zero overhead abstractions" or "abstraction costs" This is another odd term. In Rust, the term used instead is "zero-cost abstractions," referring to cases where the compiler can generate instructions for higher-level code just as efficiently. > So the main practical benefit of a JIT isn't that at all, but that it can do the "mother of all optimisations" - inlining - far more aggressively. I’ll reiterate that I disagree with this: inlining is performed very efficiently during monomorphization. And monomorphization is used very frequently in Rust. > After more than 25 years of working with C++ I don't have much experience with C++; I mostly use Rust. I can only assume that the C++ development experience is far worse than Rust - especially when trying to write software that is both reliable and fast. This may be particularly relevant to older C++. So, a Dog, a Cat, and an Abstract Mammal walk into a bar... https://godbolt.org/z/dEsW1sfM8 https://godbolt.org/z/dEsW1sfM8 I didn't want to do this, but I went ahead and created a small example showing that monomorphization and inlining work remarkably well. (Obviously, this example does not address memory management)
- pron 18d ago> For example, when developing two modules (possibly even by different teams), all they know about each other is a lean interface, without any implementation details. However, the compiler might know everything. You're talking about abstraction at the code level; I'm talking about abstraction at the language level. A virtual call means "the implementation is unknowable here", and it is, indeed, rarely knowable to an AOT compiler. > This is another odd term. In Rust, the term used instead is "zero-cost abstractions," referring to cases where the compiler can generate instructions for higher-level code just as efficiently. Rust took that term from C++ (and it had slightly different ones over the years). What it means that the language offers different mechanisms - chosen statically - with different abstraction levels (i.e. different generality) and different costs, some of which are zero, but often similar or identical-looking code at the use site, because the mechanism choice depends on some non-local information, typically associated with the type. I call it "writes like a low-level language, reads like a high-level one". This is different from C (or Zig), which usually makes the selected mechanism explicit at the use site, or from Java, which chooses the cheapest applicable mechanism at every use-site for a single general construct. The problem is that, because the mechanims is chosen statically, you need to choose the cheapest applicable mechanism, usually virally, yourself, and that over time this gets harder or things drift toward the more general and costly mechanisms. That's what Java tried to solve, but there are, of course, tradeoffs. The obvious one (which has solutions) is warmup time, because the compiler needs to wait to learn what optimisations can be applied even if they're unprovable, e.g. to learn that a polymorphic application is actually monomorphic in practice at a particular call-site (the solution is to cache the optimised machine code from one run to the next). The more fundamental tradeoffs are 1. you're not guaranteed which mechanism will be chosen, 2. there can be a bad, though amortised, worst-case due to deoptimisation (this is what happens when the compiler optimises too aggressively and then finds out it was wrong, e.g. it inlined a virtual call under the assumption it's the only target at the use site, but after a while, another target appears (in Rust/C++, you'll always pay the higher price, but there's no point at which deoptimisation occurs), and 3. you need an FFI layer, as you can't take the machine address of a compiled subroutine (as it may be re-compiled multiple times). Tradeoffs 2 and 3 are the main reasons low-level languages don't do this optimisation, and 3 is particularly important. Low-level languages are designed, first and foremost, to be low level. To do its sophisticated optimisations, Java needs to move around pointers to both code and data, which requires a clear FFI layer between Java code and anything external. Having such an FFI layer in a low-level language (and I'm not talking about Rust/C++'s thin extern FFI) defeats the very purpose of a low-level language, which is to talk directly to the hardware and OS. That is the chief goal of all low-level languages, and they sacrifice everything for it. Not only safety (Rust's unsafe is used relatively pervasively) but also performance. > I can only assume that the C++ development experience is far worse than Rust Actually, the experience in the two languages is remarkably similar, and not by accident. Rust certainly improves some details, but the overall experience "in the large" is very close. But note that the performance problem is not because of "zero cost abstractions" but because of the low-levelness and focus on the worst-case. Even in Zig, which tries hard to avoid zero cost abstractions to keep use sites explicit, the choice between a specific-and-cheap and a general-and-expensive mechanism means that for best performance you need to pick a less general mechanism, and that gets trickier and trickier to maintain as the program evolves over the years, and especially if it's large. > monomorphization and inlining work remarkably well. Of course it does, which is why the optimising JIT was invented: to make it work more broadly! This wasn't done just on principle, but to solve a very real problem. What we used to do in C++ is architect a solution and write code that monomorphises in all the right places - because that's what one does - and the result was good and fast. And then, five years later, we had to add some feature and were faced with the choice of either undoing some core optimisation or re-architecting some 10,000 LOC. The problems didn't arise when first writing the program, when everything was known. It arose when some change - that hadn't been foreseen when the program was first written - had to be done. Java didn't make the first step substantially cheaper; it made all the following work - five, ten, fifteen years down the line - substantially cheaper. An important caveat is that HotSpot currently misses many auto-specialisation opportunities that it could take advantage of, but that's one of the things that make working on such a cutting-edge compiler so interesting :) The problem, as always, isn't just the work required, but also determining which optimisations actually make a difference in real programs (and not just in specific benchmarks). Of course, now there's this hypothesis that AI could do this costly rearchitecting for you, even in large programs, but it doesn't do it well (at all!) today, and I think that when we get to a point where it can do it well, it will also be smart enough to do it in machine code directly (or at least in C), at which point all programming languages will be over. What I don't think is likely is that AI will be able to do extremely complex semantics-preserving large-scale transformations correctly, yet still need the help of a sophisticated compiler for much more local transformations and far simpler correctness checks.