32 ms·
Java JIT vs. Java AOT vs. Go for Small, Short-Lived Processes
- robfig 7y agoIt's hard to square these articles with the reality I see on the ground: our baseline memory usage for common types of Java service is 1 GB, vs 50 MB for Go. We do have a few mammoth servers at the top end in both languages though (e.g. 75 GB heaps) The deploy JARs have 100+ MB of class files, so perhaps it's a function of all the dependencies that you "need" for an enterprise Java program and not something more fundamental. These blog posts also present AOT as if its just another option you can toggle, while my impression is that it's incompatible with common Java libraries and requires swapping / maintaining a parallel compiler in the build toolchain, configuring separate rules to build it AOT, etc. I don't have actual experience with it though so I could be missing something.
- mpweiher 7y ago> It's hard to square these articles with the reality I see on the ground This is not a new phenomenon: back in 2000, the IBM System's Journal had a special issue on Java. Most of it was filled with articles detailing how amazingly awesome performance now was and getting better all the time. The last article in that issue was from IBM's San Francisco Project, the only one with real-world usage. Which reported much more, er, "mixed" results. More accurately: atrocious performance on their real-world tasks, and a lot of extra engineering effort to get it to work at least half reasonably. (I write a little more about this in "Jittterdämmerung" https://blog.metaobject.com/2015/10/jitterdammerung.html https://blog.metaobject.com/2015/10/jitterdammerung.html)
- evacchi 7y agoit's not "free", but there are frameworks that may help with that; e.g. Quarkus [1] provides a common EE/MicroProfile stack that does a lot of AoT codegen to get faster boot times and lower memory footprint; it is also compatible with GraalVM AoT, for even lower footprint and faster bootup. Compile times are still large, but you get the benefit of a familiar stack. Notice that, however, if your workload is throughput-intensive, you may still want to go with JVM mode (which is still pretty fast because of codegen) (full disclosure: I work at Red Hat and I have contributed to the project) [1] https://quarkus.io/ https://quarkus.io/
- vips7L 7y agoI've written a small project in Quarkus and I have enjoyed it. My main problems are the windows support still isn't there yet and the support for graal 19.3 (the one that supports java 11) isn't done yet either.
- blinkingled 7y ago> baseline memory usage for common types of Java service is 1 GB, vs 50 MB for Go. There's nothing inherent to the JVM that'd need 1GB of memory footprint - the jars are compressed and loaded on demand so unless your server needs all of them instantiated at once that doesn't explain the memory usage. Typically the region where class metadata are stored is the PermGen and it is not more than 256M for even the most heavyweight programs. My guess is that in your case the Xms (min heap) is set to 1Gb and so that's what you see as baseline. OpenJ9 is designed for lower memory usage than the hotspot VM and also supports AOT by the way if memory usage is your concern. Not disputing that the JVM needs more memory than anything else because it comes with a lot of features and it is very easy to use more but many J2EE apps can be made to work with a Gig and you do get a lot of features for that. And memory is cheap so you can decide what to optimize for.
- shizcakes 7y agoYou're almost certainly right about Xms being set to 1GB. However, even if you can experimentally set it lower, the first time the JVM app hits GC pressure, the first thing anyone is going to try is bumping that back up over 1GB to give it some breathing room. Memory may be "cheap", but wasting 950MB of memory per process because the GC might flip out at the wrong time isn't cheap when you multiply it out to many processes. Also, I find the claim that the JVM starts and is running in 90ms very dubious. I would like to see an all-in timer, something like run each in a container, and see which completes first from container startup to shutdown.
- blinkingled 7y ago> first time the JVM app hits GC pressure, the first thing anyone is going to try is bumping that back up over 1GB to give it some breathing room. Yeah but in that case your app legitimately needs that much. Not the JVM. Secondly unless your app is extremely performance sensitive setting Xms to lower value doesn't make the GC flip out - it's not as if the GC is going to collect repeatedly to stay within your lower bound - on the contrary it will expand the heap upwards until Xmx limit. Sure there will be a cost to expand the heap above Xms up to Xmx but it is not at all significant due to how clever the GC is.
- shantly 7y agoIt’s been the case ever since I can remember that “Java’s basically as efficient as native, it’s super fast” but all the actual java software I encounter is a slowish, bloated memory hog. I don’t know why, but that’s how it is.
- yumaikas 7y agoA big part of that is that this benchmark is not representative of what general software is like, not in Go, not in Java. For one thing, most software isn't churning on an array of integers. It's allocating and managing various types of objects. I haven't written a ton of java, but I suspect that the various reflection-heavy frameworks also impose a lot of overhead into the resulting software that isn't necessarily inherently part of Java the language.
- slaymaker1907 7y agoReflection actually isn’t too inefficient, especially compared to using dynamic languages. I think a lot of the inefficiency is a bunch of small inefficiencies which combine together to make it very heavyweight.
- chrisseaton 7y ago> Reflection actually isn’t too inefficient My understanding is that it's completely uncached. It couldn't be any less efficient.
- kasperni 7y agoThe overhead of, for example, invoking a method via reflection is measured in single-digit nanoseconds. Still not as fast as direct invocation, but unlikely to be a bottleneck unless you perform 10-of-millions of them per second.
- LgWoodenBadger 7y agoActually finding the method via reflection is the time consumer. Once you have the method reference, actually invoking it isn’t any different than a regular method.
- walkingolof 7y agoIsn't one difference also between the JVM heap and the Go heap that the former allocates upfront, while Go allocates when it needs more?
- sorokod 7y agoYou may be interested in this: https://blog.golang.org/ismmkeynote https://blog.golang.org/ismmkeynote
- walkingolof 7y agoThanks, seems like MaxHeap Is yet to be released feature of Go.
- nradov 7y agoIf you want to square it with what you see on the ground then you'll have to take a heap dump and analyze where the memory is actually being used. (The number of class files is unlikely to be a factor.)
- royjacobs 7y agoYou are comparing apples to oranges, then. A modern Java framework like Spring Boot, Micronaut or Quarkus can start up with a 20mb heap and will require a Docker container or jar in the range of tens of megabytes. Now, clearly to do more complex processing you'll want to give it more than tens of megabytes of heap, but if you are running Java applications that needs hundreds of megs of dependencies and 1gb of memory then you're probably running some legacy application server.
- geodel 7y ago> These blog posts also present AOT as if its just another option you can toggle, while my impression is that it's incompatible with common Java libraries I agree. I have experimented to make one of my smaller application (~20 Java classes) to Graal native image. After 4-5 days of trying I gave up. Most of the commonly used libraries like JDBC drivers, Kafka, log4j etc work in totally opposite way that Graal Native image like to work. The way I see it is J2EE all over again. Various vendors will find overall complexity of native images for practical or business Java apps as consulting or business opportunity. We can already see Quarkus, Micronaut and others are trying to offer native images for their framework. This would mean only a set of libraries would be compatible to native image in practical sense.
- Twirrim 7y agoThe company I worked for a couple of jobs ago used to host approximately 130-140 java based web applications, all using Spring, on just 4 servers (primarily hosted on 2 of those) With only a few exceptions the java webapp was configured for a 256MB heap. Devs followed good practices to avoid creating excess garbage, and kept the number of libraries down to a minimum. They also struck a good balance between doing transformations in-memory, and on the database side. side note: There was _one_ web app that we had to host that wasn't developed in-house. It was absolutely god-awful, badly scaling, websphere based crap (I have no idea if Websphere is crap, but that application sure was bloated and underperforming). Websphere was irritating to deal with. What they did with it was worse. We eventually got the source code from them so we could fix it up, after a bunch of legal wrangles. They had 4 different custom created ORMs in it, each clearly developed by a different dev. Then for reports they never used the ORM, instead they'd make a "one query spits out the final report" query. Those queries would often be 30 - 40 subselects deep. `SELECT thing FROM foo WHERE bar in (SELECT bars IN monkey WHERE ....` and so on down the line.
- an_d_rew 7y agoIt is absolutely possible to write amazingly performant Java / JVM applications, and absolutely possible to write horrible blobby slow Go applications. But given most of the standard corporate devs that I've worked with, it is infinitely easier to write crappy Java and not-bad Go than the other way around. Sadly.
- ivan_gammel 7y agoIt's also way easier to find a Java developer. Go is still in early adoption stage and people don't start programming on Go because they've decided to change their job from marketing to IT. As soon as any platform becomes popular the volume of crappy code increases dramatically.
- andrewprock 7y agoJust wait until enterprise Go is released.
- edem 7y agoI don't know what are you using but I've built apps using Spring (known as a resource hug) which take up 10MB of memory and start up super fast. You only need some fine tuning to get these numbers. Alternatively you can also use things like Quarkus: https://quarkus.io/ https://quarkus.io/ which are hilariously cheap to operate and can work with Lambda as well.
- imtringued 7y agoIt's a problem that is inherent in the way the JVM's GCs work. They never release memory even if you have low heap usage. I have a micronaut based microservice. Visual VM reports less than 30MB memory usage. Sometimes I get request spikes and it goes to 200MB. It will allocate 200MB from the OS and once the spike is gone it will hang onto the 200MB. There is also a fixed overhead. Even if I set the heap maximum to 30MB at the cost of running out of memory during spikes the total memory used by the JVM is still around 130MB. If you want to run 10 microservices that each do very little then you will need anywhere from 1.3GB to 3.3GB (as soon as you have one traffic spike and then never again). With go you could probably get away with just 300MB and during temporary traffic spikes your maximum will be 2.3GB at most but it is unlikely that all your microservices will spike at the same time so you will need significantly less than 2.3GB.
- miskin 7y agoIt depends on Java version that you use. Recent GC returns unused memory back to OS [1] and class data sharing can speed up startup time of your microservices and save memory as shared classes are loaded only once for all JVMs [2] [1] https://openjdk.java.net/jeps/346 https://openjdk.java.net/jeps/346 [2] https://openjdk.java.net/jeps/310 https://openjdk.java.net/jeps/310
- sreque 7y agoThe problem is enterprise Java and the Java community, not the jvm. If people wrote more god-like code in Java, we'd have much leaner Java processes. Instead, we have heavyweight, performance-insensitive frameworks like spring and hibernate used everywhere when they shouldn't be.
- ignoramous 7y ago> These blog posts also present AOT as if its just another option you can toggle, while my impression is that it's incompatible with common Java libraries and requires swapping / maintaining a parallel compiler in the build toolchain, configuring separate rules to build it AOT, etc. No, the blog post doesn't, and I quote: "Java AOT comes at a price. Huge compilation times (may slow-down your CI/CD pipeline). Very limited reflection API, limiting the usage of several frameworks or requiring of extra, complex, configuration: JPA, Spring... (This aspect will be treated in a future blog post)."
- moomin 7y agoI mean, that’s all well and good, but the article is about short-lived processes and you seem to be talking “enterprise” long-lived processes.
- nine_k 7y agoAt one of my recent employers, we wondered why some of our services consume whopping 200 MB when a normal service would run on 50 MB. A service requiring 1-2 GB would be considered a behemoth. Unless your code is extremely complicated, or you are running a Jack-of-all-trades monolith, the only way to reasonably consume gobs of RAM is to keep an in-memory database / cache. This trades RAM consumption for speed. This is what e.g. IntelliJ does. Running a heap larger than you need slows you down in the long term, because GC over a large heap consumes more time and CPU.
- c048 7y agoLook into Micronauts and GraalVM. Simple apps have shown to have startup times of sub 20 ms and memory footprints als low as 18 mb. It's simply not fair of a comparison to GO when you drag in heavy frameworks like Spring, Hibernate, etc... . I'd also like to add that even without GraalVM and Micronauts, an app that requires 1 Gig of memory seems to be indicative of its architecture and not the choice of language.
- nobleach 7y agoMicronaut and Quarkus are fantastic! Those startup times are exactly what you're looking for if you want to run a cloud function or lambda. I will say, that for a standard microservice, where startup time is not as important, I'd not compile to GraalVM. The cpu/memory overhead to create a GraalVM image is pretty intense. Furthermore, GraalVM doesn't optimize as it warms up like HotSpot does. That need may depend on your use case. I'm currently running Micronaut services in production that take around 5 seconds to start on JDK 1.8. The Spring Boot equivalent..... well.... 10 times that amount of time maybe?
- mikece 7y agoI'm curious how Dart2Native would fare in this test, if it would be more or less the same as Go or if it would be more efficient.
- trimbo 7y agoIs that just curiosity or are you using it? What kinds of things are people using Dart2Native for?
- erokar 7y agoI'm curious about this too. Dart has an interesting runtime since it supports both quick JIT compilation and AOT to machine code.
- mikece 7y agoWith Flutter it's fully AOT compiled to binary ARM code and is a compelling choice for IoT applications as well.
- ape4 7y agoFor a short lived Java process you can use the no-op garbage collector (ie don't collect). http://openjdk.java.net/jeps/318 http://openjdk.java.net/jeps/318
- MaxBarraclough 7y agoThat would save the trouble of initializing a garbage collector that isn't going to be used. Is that a significant saving?
- jillesvangurp 7y agoInitializing a GC has almost no overhead (compared to e.g. loading some classes). The only reason to introduce this is to intentionally avoid overhead for having it actually reclaim memory under the assumption that it is in any case not going to ask for more than there is. This does not make sense for long running servers but could make sense for short running things like lambda functions, command line stuff etc. Of course if this overhead is substantial that probably means this is not a good solution since you apparently have a lot of memory allocation happening.
- stygiansonic 7y ago“Last-drop throughput improvements. Even for non-allocating workloads, the choice of GC means choosing the set of GC barriers that the workload has to use, even if no GC cycle actually happens. All OpenJDK GCs are generational (with the notable exceptions of non-mainline Shenandoah and ZGC), and they emit at least one reference write barrier. Avoiding this barrier can bring the last bit of throughput improvement. There are locality caveats to this, see below.” From: https://openjdk.java.net/jeps/318 https://openjdk.java.net/jeps/318
- MaxBarraclough 7y agoI'd missed that, thanks.
- chrisseaton 7y ago> That would save the trouble of initializing a garbage collector that isn't going to be used. It also saves evacuating thread-local allocation spaces and running all the barriers.
- deleted 7y ago[deleted]
- skywhopper 7y agoInteresting numbers. I have been out of the Java world for a few years and I’m unfamiliar with GraalVM but I’m curious how compatible it is with Oracle Java or OpenJDK. Of course for small scale stuff like a QuickSort implementation, JVM starts are fast-ish, but for a nontrivial service, library load times during boot can balloon quickly depending on your build discipline.
- evacchi 7y agoincredibly compatible, albeit with some limitations. The compiler is quite aggressive, so e.g. you can't do reflection on "any" class, you just have to tell the compiler what you are going to use at runtime (so called "closed-world assumption"). Also most initialization code is forced to run at "compile-time" to keep boot time low. Pretty incredible piece of code.
- mcguire 7y agoWhat's the license on it? I'm looking at a project that makes "The main benefits of doing so is to enable polyglot applications (e.g., use Java, R, or Python libraries)" sound attractive.
- ddtaylor 7y agoFor anyone looking to avoid attacks because of no SSL: https://archive.is/Rzdko https://archive.is/Rzdko
- harikb 7y agoWhen I was looking for my first car in early 90s, I knew nothing about cars or brands. One thing I noticed was that all the TV ads for most of the cars would say "more room than a Camry". I knew what I needed to buy. If Go doesn't survive another decade, I would still be happy about what it triggered.
- michaelcampbell 7y agoSince the article is about "short lived processes", I'm having a hard time caring about the memory use as a first order issue. <shrug>
- AnimalMuppet 7y agoDepends on how many of them you have running at any one instant.
- firethief 7y agoIf you have short-lived tasks popping off at that rate, the JVM looks some orders of magnitude better if you don't spin up a fresh process for every task.
- haolez 7y agoI wonder if AOT would speed up Groovy as well. My intuition is that it doesn't matter, since the runtime will get bundled and execution time will be the same.
- imtringued 7y agoThat's exactly what happens. When you compile Groovy with GraalVM it will generate a fallback image that still requires a full JVM.
- zestyping 7y ago100 ms is nowhere near what I'd call "negligible" for a process that might only live a second or two! python -c print 'hello' starts up and shuts down an entire Python interpreter in less than 50 ms on my machine, whereas the equivalent Java program never takes less than 120 ms. That seems pretty sad for a language that's had at least an order of magnitude more resources thrown at it.
- pron 7y ago> whereas the equivalent Java program never takes less than 120 ms Then you must be using an old version of Java.
- Groxx 7y agoGiven that hardware varies, and the article was showing 80-90ms floors: 120ms floor seems entirely reasonable to me.
- pron 7y agoThe article wasn't using a current Java, either. Current numbers are under 40ms for Hello, World: https://cl4es.github.io/2019/11/20/OpenJDK-Startup-Update.html https://cl4es.github.io/2019/11/20/OpenJDK-Startup-Update.ht...
- ptx 7y agoThat's certainly good news and the renewed focus on start-up performance is very welcome. I wonder what the numbers look like for small Kotlin programs. (Although I see now that the comparison disabled CDS on Java 8, so the actual improvement is perhaps not as large as it seemed.)
- ivan_gammel 7y agoIn one particular case when a process will live for a second or two, it's indeed a visible delay. But it's just one of many scenarios and it's definitely not the top priority one for applying all those engineering resources. When you build a server application, you just don't care about startup time much - in situations where it matters, green-blue deployment will do a better job to reduce downtime (and, anyway, it's not always the platform startup which contributes to the delay before full availability).
- karmakaze 7y agoI've run small services built with Java, Go and Crystal to achieve good startup and throughput performance while minimizing memory usage. My experience with Crystal is limited but has been positive thus far. The sweet spot for me is using Java with OpenJ9 which has very fast startup time while sacrificing only a bit of top-end throughput. I would only choose AOT if packaging/deployment were issues with the JIT approach. In the case of Go, AOT is practically free but I prefer not being limited to array/slice, map, and channel generics.
- mister_hn 7y agoThe question is: if you compile Java to native code, what's the purpose of using Java instead of Go or Rust or C++?
- eternalban 7y agoLanguage semantics; libraries; runtime; tooling; talent pool; ... [p.s.: point is not that X is better than Y.]
- mister_hn 7y agowhat? you forgot tooling and semantics of modern C++. I found C++ tooling far superior to Java ones
- chrisseaton 7y agoExisting code. Existing expertise. Existing tooling.
- xvilka 7y agoYou can convert it to Kotlin and just use LLVM-based native[1] compilation that is even more complex efficient. It's completely automated, and gives you the better and more modern language without much hassle. [1] https://kotlinlang.org/docs/reference/native-overview.html https://kotlinlang.org/docs/reference/native-overview.html
- 7y ago
- rijoja 7y agoWouldn't the JIT:ed code be optimized for the specific CPU, whereas the AOT perhaps can not make use of certain instructions? Also, when making a comparison off the different models one must keep in mind that the software that processes the code have different characteristics.
- mcguire 7y agoThat's what people keep telling me.
- deleted 7y ago[deleted]
- deleted 7y ago[deleted]
- mdasen 7y agoJIT'd code isn't just about specific CPUs, but about optimizing hot paths. For example, you don't want to inline a commonly used function all over the place, but if there's one area that calls it 10,000 times per second while the others call it once a minute at most, you can re-write the program at runtime having observed that hot code path. To get an intuition for JITs: there are pieces of your code that you know will always be run in a certain way (or mostly run in that way), but it isn't provable at compile time that it will always be that way. JITs can notice that pattern and optimize it (or even provide optimizations for common, but not exclusive paths).
- correct_horse 7y agoGo has no proper solution to garbage collector ballast, a hack which businesses you've heard of are using in the wild. see https://blog.twitch.tv/en/2019/04/10/go-memory-ballast-how-i-learnt-to-stop-worrying-and-love-the-heap-26c2462549a2/ https://blog.twitch.tv/en/2019/04/10/go-memory-ballast-how-i... and https://github.com/golang/go/issues/23044 https://github.com/golang/go/issues/23044. A golang team member's reasons for not adding a minimum heap size include that it would require more testing each release, and that they might want to have a max heap size hint instead. I posted a comment similar to this one recently, but it seems more relevant here.
- Groxx 7y agoI've noticed similar things while profiling my tests - making a large static allocation or bumping GOGC=1000 can cause them to run more than 2x faster (my favorite took a 12 second test suite and dropped it to about 2.5 seconds). So much time is spent assisting the GC to keep memory at like an 8-12mb range, as if that small of a heap was somehow the most important thing Go could be doing with the CPU.
- sreque 7y agoThis article reiterates the belief that jvm start-up time is slow is somewhat a myth. When. I measured it years ago, I found jvm start-up time to be roughly equivalent to node.js. What makes start-up time bad in any interpreted language, including Java, python, and JavaScript, is code-loading time. This time is 0(n), where n is the total size of your app, including transitive dependencies. It takes time to load, parse, and validate non-native code. This time far dwarfs any vm start-up time. As an experiment, write a hello world node.js app and time it's execution. Then add an import statement on the aws sdk. Don't actually use it. Just import it! When I last measured it, this caused start-up time to go from 30 Ms to something like 300 Ms. The extra time mostly comes from loading code. For a native app, the binary itself just gets mmapped into memory. Shared library loading is more expensive, but not much, and way less than loading source code or byte code. The tldr is if you want a fast starting non-native app, you have to shrink the transitive dependency closure your app loads to do it's job. This is easy for toy benchmark apps but can be harder for real apps. It also goes against the philosophy of most devs to rely on third party libraries for everything. For instance, If you care about start-up time, it may be worth re implementing that function you'd normally get from guava or apache commons. You can alternatively use a tool like proguard to shrink your dependency closure.
- e12e 7y agoInteresting write-up - would be nice to see not just the quicksort code, but the harness/scripts used for benchmarking. As far as I can tell it's not included?
- gok 7y ago"3 Bad Tools for a Job, which is least bad?"
- ncmncm 7y agoThis. With a big enough hammer, every screw is a nail.
- CountHackulus 7y agoI'd kill for a comparison with the IBM J9 AOT flag. It essentially just caches jitted code for startup. Also if startup time is super important than you can try -Xquickstart.
- xvilka 7y agoIt would be nice to have some automated tools to convert from Java to Go or Rust. Something like c2rust [1], but for Java. There exist[2] some kind of automation, buts it's too basic to be practical. [1] https://GitHub.com/immunant/c2rust https://GitHub.com/immunant/c2rust [2] https://github.com/aschoerk/converter-page https://github.com/aschoerk/converter-page
- AtlasBarfed 7y agoWhy can't the JVM preload on startup the majority of the core runtime as a shared library architecture in a sort of super-permanent generation, and then the JVMs piggyback on that? I remember solaris boxes had little of the startup cost and someone told me they preloaded the java runtime.
- thu2111 7y agoIt does. That's called AppCDS and is a relatively recent feature, it's also not on by default buy they're working on making it be used automatically.