6 ms·
The memory mapping trick they use on x86 to avoid masking only works up to the maximum 48 bits of addressable virtual memory, so there is much less than 22 free
by obl 8y ago
The memory mapping trick they use on x86 to avoid masking only works up to the maximum 48 bits of addressable virtual memory, so there is much less than 22 free bits in that case.
It's also not quite free since it takes up TLB space.
- DSMan195276 8y agoAgreed - I'd go farther to say it only works for addresses in the userspace half, so on Linux you lose an extra bit and only get 47 (I don't know how the mappings are setup for Windows/OSX, sorry). It also might be worth pointing out that you now need 16x the page mappings (Since every mapping needs 15 duplicates for the possible flag states) - I don't know how Java does it's memory management but if it does lots of small mappings (Which I'm guessing it does not) then that could be a concern. And like you mentioned with the TLB, it takes up space and it also slows things down a bit since if the flags change the entry in the TLB will no longer be used (since the address is now different) and the MMU will have to walk the page-table again. In practice, I would wager the performance concerns aren't huge, but that's mostly just a guess. I'd personally be interested in seeing a comparison to masking to see just how much slower masking would be, but obviously it's not like they can just flip a switch to use masking.
- HohPum1l 8y ago> It also might be worth pointing out that you now need 16x the page mappings It doesn't. The article mentions it's only 3 mappings since those bits are colors, not arbitrary combinations of flags. > I don't know how Java does it's memory management but if it does lots of small mappings (Which I'm guessing it does not) then that could be a concern. openjdk generally uses large contiguous mappings but it may punch holes in the middle of the heap if it's configured to yield back memory to the OS. But applications that dynamically shrink and expand their heaps are not necessarily those that are concerned about the last quantum of page table overhead. > And like you mentioned with the TLB On the other hand it does support huge pages to mitigate costs of TLB entries.
- pitaj 8y ago> It also might be worth pointing out that you now need 16x the page mappings (Since every mapping needs 15 duplicates for the possible flag states) Nope. They went into this in the article: > Since by design only one of remap, mark0 and mark1 can be 1 at any point in time, it’s possible to do this with three mappings. There’s a nice diagram[1] in the ZGC source for this. [1]: http://hg.openjdk.java.net/zgc/zgc/file/59c07aef65ac/src/hotspot/os_cpu/linux_x86/zGlobals_linux_x86.hpp#l39 http://hg.openjdk.java.net/zgc/zgc/file/59c07aef65ac/src/hot... That might not be totally current, as it doesn't cover the finalizable flag, but if it works the same, that would only be four mappings. If it works differently, then it would be a maximum of 6 mappings. Not 16.
- shaftway 8y agoAnd it looks like you only need 3 mappings for the entire space. It's not like you need one per object in memory. Unless I misunderstood the gist of the article.
- deleted 8y ago[deleted]
- adrianmonk 8y agoI agree that it definitely can't be completely without cost. But I'll speculate that they may have arranged for the don't-care bits in the pointers to all be 0 when GC is not running and doesn't need the bits. If so, that could mitigate the TLB waste during periods when GC isn't running. In other words, just because the alternate mappings exist doesn't mean the pointers always have values that actually use them.
- joemag 8y agoMany Java applications will gladly take 10%, or even higher, performance hit if that means more predictable GC times. If you own a latency sensitive application, then it's common for tail latencies to be dominated by GC times. Bringing tail latencies closer to median would be a huge win for those app, even if the median itself moved.
- gct 8y agox86 caches uses virtual address for indexing (physical bits for tag), so this'll increase cache pressure too.
- titzer 8y agoIt takes more TLB entries, but not physical cache space. Virtual indexing just makes use of the page offset being the same between virtual/physical mappings in order to select the cache set before mmu translation is available.