8 ms·
> Thank to runtime information it can occasionally exceed performance of statically compiled language. Interestingly in 30 years I have not once heard of a cas
by dialamac 6y ago
> Thank to runtime information it can occasionally exceed performance of statically compiled language.
Interestingly in 30 years I have not once heard of a case where this theoretical benefit has manifested as a clear advantage in any real world application when looking at the system as a whole... amdahls law and all that.
You can always hand tune the 1-10% hotspots for reasonable cost most of the time, and even static tools can do PGO which generally gets you where JIT would anyway.
- chrisseaton 6y ago> in 30 years I have not once heard of a case where this theoretical benefit has manifested as a clear advantage in any real world application Today you make make a very direct empirical comparison to see this - using the Graal compiler. This lets you compile exactly the same Java code either ahead-of-time or just-in-time, but using the same compiler logic except for the runtime information available when running just-in-time. The just-in-time code is (ignoring startup and warmup time) in my experience always faster, due to the extra runtime information.
- darksaints 6y agoHave you tried graal's PGO?
- chrisseaton 6y agoNo unfortunately - it's a closed-source Enterprise feature I believe. But the current differential is pretty large and nobody is shouting loud that they can fix it using the PGO that I've heard. And what will the PGO determine that a JIT can't also do?
- darksaints 6y agoA number of things actually. For one, most JIT implementations only optimize once, instead of continuously. The result is machine code that is optimized for the sorts of things done at startup, as opposed to steady state operation. For example, I have a Play app that takes a minute to start up. The JIT does a great job of optimizing the code that is called during the setup process, but the API code itself doesn't get much optimization. With PGO, I can get a more representative profiling dataset, allowing the JIT to see actual production loads instead of startup loads. And for CLI apps, PGO code starts up fast and never slows down to profile or optimize.
- chrisseaton 6y ago> For one, most JIT implementations only optimize once, instead of continuously. The result is machine code that is optimized for the sorts of things done at startup Funny you should mention that - the author of this blog post has another post on fixing that problem for one specific (but very practical) case where we want to disregard some profiling information from the startup phase because it pollutes the genuine profiling information. https://engineering.shopify.com/blogs/engineering/optimizing-ruby-lazy-initialization-in-truffleruby-with-deoptimization https://engineering.shopify.com/blogs/engineering/optimizing...
- saagarjha 6y agoMost JIT compilers will go after any code that shows up as hot, regardless of when it executes. If your API code is that, even a minute after startup, it should really be getting optimized…
- darksaints 6y agoThink of an use case where you have an API whose mode behavior is to do nothing, waiting for a call... but commonly gets calls that are very compute intense for short bursts of time. With the most common type of JITs, which profile once and compile once, I'm going to get code that is optimized for startup and initialization. If I have an "advanced" JIT, which is constantly deoptimizing and reoptimizing for whatever it sees as the hottest path for some arbitrarily chosen snapshot of time, I'm going to see my compute-intensive code slowed down so that it can optimize and compile it every time that endpoint is called, but then subsequently deoptimized while it is sitting around waiting for something to do, ensuring that I have to go through the same process the next time it is called. You can actually see a lot of situations where this regime could be even worse than a naive startup-based single optimization, which is why it is actually not that common outside of dynamically typed languages. With PGO, I can select a profile snapshot during a stress test, and get heavily optimized code specifically for the things that actually stress the server. And it will stay optimized for that use case.
- tsimionescu 6y ago
- ptx 6y agoIsn't this because Java is designed for JIT compilation, or at least not designed (with the appropriate tweaking knobs) for AOT compilation? Languages built with AOT compilation in mind (e.g. Rust or Nim) usually give you lots of ways make choices at compile time and give hints to the AOT compiler that the JIT compiler would instead try to infer at runtime in Java. But by infering these things at runtime intead, maybe the JIT approach makes it easier to get fast code in those cases where you (as an application developer) don't want to put a lot of effort into optimization?
- chrisseaton 6y agoYes... but then we're not comparing JIT/AOT anymore - we're comparing different language designs.
- pjmlp 6y agoJava has had commercial implementations of AOT compilers since the early 2000. Most compilers for embedded systems have always offered that option, and in what concerns enterprise JVMs, JIT compilers have had the capability to cache JIT code and PGO data between runs. Both options that have come now to OpenJDK, OpenJ9 and Graal. Android also learned the hard way that changing to pure AOT did not achieve the performance improvements that they expected, while compilation on device achieved C++ compile times when it was time to update all apps, hence the multi-tier interpreter/JIT/AOT with PGO introduced in Android 7. The main problem of AOT compilation with PGO, is that first of all one needs a good dataset so that the optimizations are in line with the actual behaviour in production, still doesn't work across dynamic libraries so optimizations like devirtualization are not possible, and most of the time the tooling is quite cumbersome to use.
- seanmcdirmid 6y agoPretty sure assymetrix was done in the late 90s. 1998 or so, they started out doing educational software and pivoted to doing ahead of time compilation for some weird reason. Here is a link to when it failed in 1999: https://www.cbronline.com/news/sepercede_exits_java_sells_to_instantiations/ https://www.cbronline.com/news/sepercede_exits_java_sells_to...
- dialamac 6y agobut from a higher level.. you could just reimplement the thing in C++ with gcc and the whole thing will probably perform better. Basically what I’m saying is that I’ve never seen any substantial rewrite of decent C++ code into Java perform better, even if jitting has benefits on a small scale, it’s not substantial enough to overcome other overheads in managed languages.
- saagarjha 6y agoC++ and Java have other language-level differences, though, that go beyond just "JITs are slower".
- littlestymaar 6y agoTrue, but part of that difference comes from the “we don't need to design the language for perf, the JIT will close the gap automagically” mindset though.
- saagarjha 6y agoI don't really think so for Java; most of it was designed prior to high-performance JIT-based VMs were a thing. In fact I think a lot of the advancements in JITs actually came out of work that went into making HotSpot fast.
- dialamac 6y agoSure but one of vocal pitches/hype for Java in the beginning and even smalltalk before that.. was don’t worry about the price paid for .. automatic memory management, bytecode, dynamism... the top men are working on JIT, GC research and other optimizations that would not only close, but handily overtake the gap compared to the state of the art static compiled languages of the time. (Hence exactly why Hotspot was a thing after java was clearly becoming widespread) In reality it was a bit of a mixed bag.. and to some of us that remember the hype from 30 years ago it comes across as over promising and underdelivering. That isn’t to say that the technology isn’t incredible, I don’t mean to dump on it. But overpromising is sort of the status quo for tech.
- dahfizz 6y ago>> Interestingly in 30 years I have not once heard of a case where this theoretical benefit has manifested as a clear advantage in any real world application when looking at the system as a whole. > The just-in-time code is (ignoring startup and warmup time) in my experience always faster, due to the extra runtime information. I don't think that result is surprising. The issue is that in the real world you can't ignore startup and warm up times. I've never heard the claim that JIT compiled code is slower than statically compiles code. The issue is that the extra costs associated with JIT don't outweigh its benefits.
- darksaints 6y agoI too keep hearing this repeated but without empirical support. Same goes for the idea that garbage collection can be faster than manually managed memory. In practice, PGO has been the best possible compilation regime that I have ever found.
- pjmlp 6y agoIt goes both ways actually.
- saagarjha 6y agoIt's easy to come up with cases where you can do better than a fairly reasonable programmer: think of a loop that allocates at the top and frees at the bottom with enough stuff happening in between that the allocator doesn't just hand you back the same memory from a cache. Depending on the runtime, it might be better to just bump allocate and garbage collect all that memory at the end. You can find the benchmarks that'll show it, too. In reality, with actual applications garbage collectors have "good enough" performance with many benefits such as correctness and ease-of-use, but manual memory management done by a skilled programmer will outperform it.
- dialamac 6y agoI think that’s my point. I get that JIT can produce faster code, I get that java is good enough for most things. But if you’re already in the performance critical regime there are other things you have to do where just having JIT available isn’t some magic bullet.
- kipply 6y agoA possible magic solution is to load a JIT state from a previous execution and have good manual-optimization features to be like PGO. This would give JITs best of both worlds with being able to get to a strong peak performance and also not requiring manual tuning
- littlestymaar 6y ago> Same goes for the idea that garbage collection can be faster than manually managed memory. It's really workload dependent and depend a lot of the GC involved (a pretty dumb one like Python's or Go's won't get you anything performance wise), but a copying collector can achieve allocation way faster than a regular heap allocator (the allocation can be almost as cheap as allocating on the stack). If you can't avoid boxing and your objects aren't all long-lived, you can run circles around a program not using such a GC.
- pjmlp 6y agoExcept that PGO is in the box for most production quality JITs, the PGO data is even optimized between executions, while most developers never bother with using the PGO toolchains for languages like C and C++. Even if they lose in micro-benchemarks championships, it hardly matters in most enterprise codebases.
- deleted 6y ago[deleted]
- kipply 6y agoJITs vs static compilers on a language made for static compilers is probably a contest that JITs will never win. The keys to making JITs "better" than static compilation is (by my limited knowledge) creation of languages that are made to be JIT compiled (like Java) and ensuring that tuning JITs is easier than PGO (though nothing that is currently in development seems promising to this end)
- munificent 6y agoThis is exactly right. You do see JITs winning in Java because Java's "everything is virtual" and "most for loop are interface calls to Iterator<T>" makes it very hard to statically compile efficiently. Languages designed for static compilation generally consider virtual dispatch to have a user-visible cost and don't unilaterally make users paying it without them asking for it.
- jecel 6y agoThe way I explain the unreasonable effectiveness of JITs and adaptive compilation (multiple compilers with the advanced ones focusing on the hot spots) is by contrasting constants and variables. In a long enough time frame nearly everything is variable: over decades even computer architectures and language syntax change. In a short enough time frame nearly everything is constant: in a single cycle a von Neumann computer will only change a single memory location and all the others will have a constant value. Even though an AOT compiler might have minutes or more to work on a code fragment, the code it generates has to work for a long time and on many different machines. A JIT might have milliseconds or less to work on the same code fragment but what it generates only needs to work on this exact machine and only for minutes or hours. In fact, it might even be wrong and have to be recompiled less than a second after the JIT produced it. So the code generated by the AOT must treat as variable things that the JIT can pretend are constant (with hook to recompile if it proves not to be so). This is slightly related to partial evaluation, which is a key idea in Graal.