5 ms·
The jstat output also shows that CCS is at 99.21% usage, which could support this theory. At a previous company, we operated some Scala services and ran into a
by cldellow 4y ago
The jstat output also shows that CCS is at 99.21% usage, which could support this theory.
At a previous company, we operated some Scala services and ran into an issue like this. I forget if it was triggered from a JDK update or a Scala update. Scala (at least the 2.x series) generated a lot of classes, so there was a lot of memory pressure on this part of the system.
IIRC, we increased the limit and it resolved the issue.
I feel like there's a flag you can pass to see JIT invocations, which might also help validate this as the problem. -XX:+PrintCompilation maybe?
Caveats: It's been a long time, so my memory may be faulty, and this may not apply any longer.
- xxs 4y ago>The jstat output also shows that CCS is at 99.21% usage, which could support this theory. If code cache has run out, the process effectively runs in interpreted mode. I'd wonder however how they would have so much code. Still, they should just run a profiler or any java monitoring tool.
- munificent 4y ago> I'd wonder however how they would have so much code. They probably didn't author that much code, but Scala 3 may have many language features that its compiler desugars to large amounts of generated code. (For example, when C# added support for anonymous functions, they initially did so by compiling each lambda to a generated class with a field for each local variable that the function closed over.)
- xxs 4y agoThey still need to use the functions quite a bit to trigger at c1. Normally Java doesn't compile immediately but after enough iterations. Also if scala does that for real, it'd eat the inline budget effectively killing the performance.
- toast0 4y ago> Normally Java doesn't compile immediately but after enough iterations. I've never done perf work with JVM, but ... Is that iteration count reduced over time at all? If not, and from other behavior described, it seems like this could explain the 48 hours of fine, and then spikey garbage. Compilation cache is nearly filled, some bit(s) of code get called often enough to be compiled after 48 hours, finding space in the cache takes lots of cpu, eventually something more useful is evicted, but then you have even more cache pressure because those more useful items will come back in less than 48 hours, and more things will come in as they hit their required 49,50, etc worth of iterations. If it's inexepensive to increase the cache size, seems like something reasonable to try, but the 48 hour period of stability makes testing difficult. I'm assuming the only realistic test system is production, cause that's how it usually is.
- xxs 4y ago>If it's inexepensive to increase the cache size, seems like something reasonable to try absolutely, you just never let that thing drop below 80%. My advice would be to also run C2 compilation at 100-1000 cycles (e.g. -XX:CompileThreshold=100) - it's a much slower startup (but they do have lots of CPUs), and it may not generate a good perf guided code but they will have a good idea how much code cache to dedicate.
- xendo 4y agoThey are using JDK17, which segments the cache into profiled and non-profiled code. Both at around 120MB by default, it’s not that difficult to hit.
- agilob 4y ago>The jstat output also shows that CCS is at 99.21% usage, which could support this theory. If the code cache is full, the cache sweeper will have more work to do, will run slower, and this cache is a linked list (afair), any attempt to create another C1/C2 optimised code will cause the allocator to treverse the list, try to find enough contiguous space, and fail, triggering an attempt to fragment the space. Occasionally, removing some less frequently used compiled caches. This process isn't your normal GC process. If it runs out of memory and nothing can be removed, you're at plateau of how fast code can execute, but your JVM is consuming more CPU cycles, that means, you're losing overall performance. There is no OOM error here, it all fails and slows down silently. No exceptions, no logs, nothing but wasted CPU cycles. This is one of the worst aspects of JVM to monitor and tune. I don't know of any promethues-like metric exporters that can be used here, like in any GC activity or stack/heap metrics. As I stated in another comment, try `-XX:+PrintCompilation` and `-XX:+PrintInlining` and `-XX:+LogCompilation`. When this turns out to be filled, try increasing `ReservedCodeCacheSize`. This is out of your non-heap area.
- Tsarbomb 4y agoStupid question, if you already are at the edge of max heap size for allowing compressed OOPs, can increasing ReservedCodeCacheSize kick you into 64 bit uncompressed land?
- agilob 4y agoI'm not sure if I understand the question and context, but reserved code isn't in heap, but in "non-heap area", non-heap area has many sections, so you could potentially cause OOM on non-heap somewhere, or just see another regression in performance, due to poor heap vs non-heap ratio, then you change default -XX:MaxRAMPercentage=25, to something lower like -XX:MaxRAMPercentage=22 (depends on your total available memory).