10 ms·
From Java code to Java heap (2012)
- chii 10y agovery detailed and interesting read!
- dmichulke 10y agoMost useful info: If you run the 64 bit JRE and have memory problems, use the flags - Xcompressedrefs - XX:+UseCompressedOops for a decent 35-45% reduction in memory use. The rest is well-known stuff (use primitives, use arrays with fixed size, ...) Still, the numbers are quite interesting (and the overhead quite scary)
- the8472 10y agoOn hotspot ergonomics automatically turn those options on as long as your max heap is < 32GiB. $ java -Xmx31G -XX:+PrintFlagsFinal 2>&1 | grep UseCompressed bool UseCompressedClassPointers := true {lp64_product} bool UseCompressedOops := true {lp64_product}
- nalllar 10y agoThey also can't be enabled with >32GiB heaps.
- the8472 10y agoThey can if you increase the object alignment, but that's generally not advisable and should only be done when measurements show a benefit.
- bogomipz 10y agoI think it depends on the JVM. You can use compressed refs on JRokit with > 32G heaps. Its a bit of "bit twiddling." This article explains it pretty well: https://blogs.oracle.com/jrockit/entry/understanding_compressed_refer https://blogs.oracle.com/jrockit/entry/understanding_compres...
- PDoyle 10y agoThis article is super old. Those options are now default, and the overhead is much smaller. (I think even in 2012, we already had 1-word headers.)
- stuff4ben 10y agoAs I move into other languages like Go and Swift, would love to see breakdowns of how they store objects in memory comparable to Java.
- pjmlp 10y agoIn any case better, thanks to value types support and layout control. Somehow I would like they would have focused on value types and reified generics for Java 9 instead of jigsaw and postponing them to Java 10+.
- needusername 10y ago> Somehow I would like they would have focused on value types and reified generics for Java 9 instead of jigsaw and postponing them to Java 10+. I hear you. Cost-benefit wise Jigsaw doesn't look attractive to me. To be honest I would have preferred Java 9 in the fall of 2016 without Jigsaw.
- pjmlp 10y agoJava is getting lots of new competition from languages that have proper support for value types, have layout control and have AOT available for free on their reference compilers. On the desktop it already lost to .NET, Qt and HTML 5. On the mobile it might reign in Android, but Android Java is not 100% Java and I don't see Google improving the language, in case Oracle stops doing it now that they had another setback. They went silent on future of JEE and IBM is now doing lots of Swift and Go support, while making J9 language agnostic. So I really don't care about jigsaw.
- Longhanks 10y agoAlso, what advantages, except existing code bases, does Java provide to Kotlin, Go, Swift, Rust? For about any use case I can imagine a better choice than Java. Also, concerning the recent events (Google vs. Oracle), I think every new project might want to triple check if they really want to use a product from someone like Oracle.
- needusername 10y agoIt has to be noted that on HotSpot the object header is only two words not three words like on J9. Basically Flags and Locks fit into one word on HotSpot whereas they use two words on J9.
- jcdavis 10y agoOnly on 32 bit IIRC, headers are 12 byte on 64 bit with CompressedOops (the default for heaps <32gb) and 16 bytes without. The bit layout of the markoop is here: http://hg.openjdk.java.net/jdk8/jdk8/hotspot/file/87ee5ee27509/src/share/vm/oops/markOop.hpp http://hg.openjdk.java.net/jdk8/jdk8/hotspot/file/87ee5ee275...
- needusername 10y agoYou are correct.
- PDoyle 10y agoJ9 has had one-word headers for years now. I think this article was already out of date in 2012 when it was published.
- needusername 10y agoThat sounds interesting. Do you have source that you can share?
- deleted 10y ago[deleted]
- dr_rust 10y agoCan someone add 2012 in the title, Most of the OpenJDK collections have been rewritten during the Java 8 timeframe, so the values are out of date :(
- geodel 10y agoI think the object layout is not changed. I just ran jol tool for HashMap and size looks similar to what is mentioned in the article.
- dr_rust 10y agoYou may be right, i don't know. Tracking a cause of OutOfMemoryError recently (we maintain several applications written in Java at my day job), i've observed that most data-structures have a size very different if empty or not.
- hyperpape 10y agoYes, because many of them allocate a zero size array when you call their constructor and then lazily add to it. But that isn't really about object layout as I think of it, so much as what the fields of a particular object are.
- jcdavis 10y agoString is one that has changed pretty significantly - count and offset have been removed, only fields are hash and value now. This means String.substring is no longer an O(1) method, but it saves 8 bytes per string which is huge, and also avoids some of the wierd corner case GC/memory issues that substring/StringBuilder caused by hanging on to the reference
- geodel 10y agoRight. Java 9 will move from char[] to byte[] for String.value which will lead to further savings.
- needusername 10y ago
- gmarx 10y agoIn my work I rarely find this sort of thing relevant. It's much more important to make sure you don't have object references lying around. I mostly do server stuff now. Maybe this kind of optimization is still important for devices?
- breischl 10y agoAs with everything performance related, it depends. I do server-side work as well, but I've seen cases where it matters. eg, somebody got overly "OO-happy" with a response object and managed to take something that should've been one object with 8 fields and instead made a graph of 8 object with a total of 12 fields (4 of them repeated), wasting ~400 bytes each IIRC. When you're creating one of those for each request and handling 200k requests/sec it adds up to a lot of memory. That means a lot of time spent in the GC, which means a lot of GC pauses, not to mention effects on memory bandwidth, locality, and processor cache usage. All else equal, using less memory is faster and more scalable than using more memory. Java programmers seem to frequently forget that object references do have a cost associated with them. Tangentially, those complex object graphs also make (de-)serialization much harder than it needs to be. Requests & responses should be as simple as possible!
- gmarx 10y agoWhat did it save you server-wise and how long did it take to identify and make the change? Also, did you discover this as part of a performance problem investigation or was it something you saw upfront and nipped in the bud before it became a problem. I agree it sounds like poor design. I instinctively simplify whatever info is going to be sent from the server.
- whack 10y agoAs someone who used to be a hardware engineer, I found Figure 1 in the first section surprising. All modern OS run processes on independent virtual memory spaces, in order to ensure that processes don't collide with one another, or with the OS itself. But if figure 1 is to be believed, the kernel shares the same address space as the process. Is this a mistake on the part of the writer? [1] http://www.cs.utexas.edu/users/witchel/372/lectures/15.VirtualMemory.pdf http://www.cs.utexas.edu/users/witchel/372/lectures/15.Virtu... [2] https://en.wikipedia.org/wiki/Virtual_address_space https://en.wikipedia.org/wiki/Virtual_address_space
- kevhito 10y agoNo, it's not a mistake. Each process has it's own virtual address space, but part of that space is typically used to map operating system code and data. The reserved part is anywhere from a quarter to a half on a 4GB 32-bit system (I have no idea on a 64bit). Those pages are marked kernel-only. The upper half to three-quarters will be different for each process. One reason for this is so that you can easily trap to the kernel (for an interrupt, a system call, an exception, etc.) without changing the page tables at all -- you just change the cpu mode bit.
- ww520 10y agoIt's pretty common to map the kernel memory address to the same fixed range in the virtual memory address space of every process. Typically the kernel portion resides on the higher range of the memory space. For the 32-bit 2/2 split setup, 0GB-2GB is reserved for user mode and 2GB-4GB for kernel. With 1/3 split, 0GB-3GB for user and 3GB-4GB for kernel. This makes it easy to work on memory shared between user mode and kernel mode code since it's the same address space. Buffer passed from user mode to kernel mode is just a matter of passing down the virtual memory address pointer, no need to copy. The kernel code accessing it just accesses the lower range of the virtual memory space. The kernel code mapped to the same fixed range in every process also makes it easy to call kernel routine from user mode. SysCall just elevates privilege to be in kernel mode and jumps to the kernel routine at the exact same address in every process. You can think of the kernel as a special library got "linked" into every process at the exact same location, along with all its data. Although user mode and kernel mode are in the same memory address space, user mode cannot access kernel mode memory. Memory pages are protected with flags, like R/W (Read/Write), U/S (User/Supervisor). A S-marked page cannot be accessed by user mode code. Protection between kernel and user mode is still in place.
- brown9-2 10y agoIt seems really silly to compare the "search/insert/delete performance" of HashSet to HashMap to ArrayList to LinkedList, since the fundamental purpose of each class is different. Not to mention that "search/insert/delete" are three separate operations with sometimes three different performance characteristics. For instance it is incorrect to state that the insert/delete performance for LinkedList is O(n) - both are constant-time.
- swsieber 10y agoI'm not sure how LinkedList insert/delete performance isn't O(n) where n is the list length - care to shed some light? Are you assuming you have a node in the linked list? That might be constant, but the standard is to remove by index, and that's O(n).
- brown9-2 10y agoAh, I was assuming foolishly that we were talking about removing or inserting only at the head or the tail. All the more reason to be more detailed in a writeup than simply "Performance: O(n)" :)
- taco_emoji 10y agoInsertion and removal will always require traversal first since the LinkedList's node class is private.
- PDoyle 10y agoIf you're interested in memory layouts in Java, I wrote a blog post recently discussing a way to make super-tight data structures in Java: https://engineering.vena.io/2016/05/09/transpose-tree/ https://engineering.vena.io/2016/05/09/transpose-tree/