9 ms·
> In a multithreaded program, a bump allocator requires locks. That kills their performance advantage. Java uses per-thread pointer bump allocators[1] > While
by natanbc 5y ago
> In a multithreaded program, a bump allocator requires locks. That kills their performance advantage.
Java uses per-thread pointer bump allocators[1]
> While Java does it as well, it doesn’t utilize this info to put objects on the stack.
Correct, but it does scalar replacement[2] which puts them in registers instead
> Why can Go run its GC concurrently and not Java? Because Go does not fix any pointers or move any objects in memory.
Most java GCs are concurrent[3], if you want super low pauses you can get those too[4][5].
Pointers can get fixed while the application is running with GC barriers
[1]: https://shipilev.net/jvm/anatomy-quarks/4-tlab-allocation/ https://shipilev.net/jvm/anatomy-quarks/4-tlab-allocation/
[2]: https://shipilev.net/jvm/anatomy-quarks/18-scalar-replacement/ https://shipilev.net/jvm/anatomy-quarks/18-scalar-replacemen...
[3]: https://shipilev.net/jvm/anatomy-quarks/3-gc-design-and-pauses/ https://shipilev.net/jvm/anatomy-quarks/3-gc-design-and-paus...
[4]: https://wiki.openjdk.java.net/display/zgc/Main https://wiki.openjdk.java.net/display/zgc/Main
[5]: https://wiki.openjdk.java.net/display/shenandoah https://wiki.openjdk.java.net/display/shenandoah
- majou 5y agoZGC[4] in particular has me excited, enough so to want to pick up a JVM language.
- the-alchemist 5y agoZGC is already available, since JDK 15, September 2020. =) https://wiki.openjdk.java.net/display/zgc/Main#Main-ChangeLog https://wiki.openjdk.java.net/display/zgc/Main#Main-ChangeLo...
- majou 5y agoMax pause times of 0.5ms is what got me really interested. It feels like a huge trade-off of GCs is almost completely gone. https://malloc.se/blog/zgc-jdk16 https://malloc.se/blog/zgc-jdk16
- silon42 5y agoThere's still a memory tradeoff, some due to GC, some due to Java (lots of runtime reflection...). Guessing 2-4x.
- masklinn 5y agoThe GC memory overhead affects all languages with a GC more advanced than refcounting. It certainly does affect Go as well.
- kaba0 5y agoAlso, throughput. But latency and throughput are almost universally opposite ends of the same axis — that’s why it’s great that Java allows for choosing a GC implementation.
- jatone 5y agoyou can fix throughput problems by adding compute resources, you can't fix latency issues. I'll always pick a GC that optimizes latency over throughput. its easier to maintain the software.
- marginalia_nu 5y agoThis is a bit of a tangent, but you can get into situations where Java's memory-overhead becomes pretty untenable. I was in a situation of having to keep track of ~1 billion short strings of a median length of maybe 7 characters. In terms of just data, that should clock in at about 10 Gb; in practice it was closer to 24 Gb. I tried going with just byte[]-instances instead, which didn't help a lot. Using long byte[]-instances and indexing those doesn't help as much as you'd think because they get sliced up into small objects behind the scenes. I ended up memory mapping blocks of memory and basically implementing my own append-only allocator.
- hashmash 5y agoThe Lilliput project aims to address this: https://wiki.openjdk.java.net/display/lilliput https://wiki.openjdk.java.net/display/lilliput
- Thaxll 5y agoZGC and Shenandoah can be slower than G1, those are not silver bullets. The fact that there is 4-5 GCs explains the situation, there is not a single GC that is better than the others. It really depends of the workload.
- geodel 5y agoIndeed. It is strange that no official JDK document puts pros/cons of GCs packaged with standard JDKs in some kinda easy-to-read table/matrix.
- kaba0 5y agoWell, unless latency is explicitly a problem with the default (G1) GC, it probably should not be changed to begin with. It is a beast of a GC with a very good balance between throughput and latency. Also, if the latter is problematic, the first thing should be to fiddle with the singular G1 knob (one should set, unless they really know what they are doing), target pause time — throughput and latency are fundamentally opposing goals. G1 by default has low enough pauses as well, but it does increase with the speed of garbage creation. But it handles even ridiculously high throughputs with graceful degradation of pause times. Here is a really great blog on various aspects of modern GCs (and don’t forget that we are at Java 17, with small but significant GC updates in each version): https://jet-start.sh/blog/2020/06/23/jdk-gc-benchmarks-rematch https://jet-start.sh/blog/2020/06/23/jdk-gc-benchmarks-remat...
- geodel 5y agoRight it is all good. As I said how difficult it would for official document to put out a comparison table. A comparison around latency/throughput/Heapsize/Cost(Mem/CPU) should be reasonable enough for developers to choose right GC for their applications.
- sam_bishop 5y agoI agree. The author seems to know quite a bit about Go and GCs, but doesn't seem to have much experience with Java. As a Java performance engineer, it sounds like he is comparing Go to how he thinks Java works based on what he's read about it.
- pjmlp 5y agoAdditionally he doesn't seem to know that much about C#, which also has advanced GC, while allowing for C++ like memory management, if needed.
- bullen 5y agoIt's odd how most people that haven't used a VM with GC are amazed by Go (no VM) and WASM (no GC) but still fail to understand that with a GC _AND_ VM you can code something that doesn't crash even if you make a big mistake! And they haven't even bothered to try it out! To use anything other than JavaSE/C# on the server you really need very good arguments!
- pjmlp 5y agoOr being amazed by Go's compile speed, when Turbo Pascal and Object Pascal compilers were already doing that in the 1980's, or finding WASM innovative when polyglot bytecodes with no GC also go back to the early 1980's, like Amsterdam Compiler Kit EM as one example among many.
- vinkelhake 5y agoThe innovation in WASM is more about getting all the major players in the browser space to agree and support it as a first-class citizen in the web stack. That polyglot bytecodes existed in the early 1980's does nothing for the web.
- bullen 5y agoWhat a perfect example, have you even tried Java? I'm guessing you are coding for the client? C++ arrogance is the problem here. About the pipe dream of WASM there are 3 problems: 1) Compile times (both WASM and the browser) 2) If you thought Applets where insecure (btw they wheren't) wait until the .js (also a VM with GC...) bindings that you are forced to go through to reach anything from WASM securely gets attention! 3) If you build for the browser you have Intel, Nvidia, Microsoft and Google (do you work there, might explain things) to deal with on Windows. You DON'T want that... use C to build for linux on ARM/RISC-V has to be one of your options and then all that work you spent on getting WASM to work is wasted. (because you won't have the performace/electricity/money for the cycles you need in the long term) Edit: Please don't replace Python (that you should probably never have touched) with Go...
- mappu 5y ago"Scalar replacement" explodes the object into its class member variables and does not construct a class object at all. That does result in the exact same `sub %esp` (that Go would do for any struct), but it is restricted to only working if every single usage of that class type is fully inlined and the class is never passed anywhere that needs it in its object form. It's worse than what Go has. Go can stack-allocate any struct and still pass pointers to it to non-inlined functions.
- pkolaczk 5y agoScalar replacement does not work even in very trivial cases: https://pkolaczk.github.io/overhead-of-optional/ https://pkolaczk.github.io/overhead-of-optional/ In all those cases, Optionals were inlined, didn't escape, yet they haven't been properly optimized out.
- deleted 5y ago[deleted]
- chrisseaton 5y agoDid you understand why?
- pkolaczk 5y agoI don't know the exact reason here, but from my experience JVMs don't seem to perform optimisations as deep as static compilers. You can see the compiler not only missed scalar replacement here, but also didn't use a conditional move to avoid branching neither it performed any loop unrolling. Maybe it is because JITs are constrained by time and resources a lot more.
- chrisseaton 5y agoIt'd be worth doing a deep dive into this - you can look at the compiler's intermediate representation to understand why it didn't make an optimisation. There may be something trivial getting in the way.
- altfredd 5y agoScalar replacements (as currently implemented in Java) does not work in real-world programs. Well-written code does not need it. Poorly written code can not trigger it, because the JIT is too dumb and isn't getting better. There is no sane test to determine whether a piece of code will be inlined in Java. In practice anything more complex than byte array is unlikely to be inlined. Even built-in ByteBuffers aren't! Meanwhile Go compiler treats Go slices just as nicely or better than arrays.
- deleted 5y ago[deleted]
- pjmlp 5y agoWhich JIT? There are plenty of them to choose from across Java implementations.
- native_samples 5y agoThe JIT is getting better. Major escape analysis upgrades are a big part of where Graal (a drop-in replacement for the HotSpot JIT) gets its performance boosts. EA definitely does work well there because Truffle depends on escape analysis and scalar replacement very heavily. GraalVM CE is better than regular HotSpot at doing it and GraalVM EE is even better again.
- altfredd 5y agoGraal effort is over 10 years old. Jaotc, — Graal's sole mainstreamed part, — has been recently removed from OpenJDK in 17th release, and Oracle says [1], that they are "considering the possibility of using C2 for ahead-of-time" 1: https://mail.openjdk.java.net/pipermail/discuss/2020-November/005634.html https://mail.openjdk.java.net/pipermail/discuss/2020-Novembe...
- kaba0 5y agoIt’s almost as if graal and openjdk are separate projects with different use cases. Also, they are not competing (in the usual meaning) since both are developed by Oracle.