3 ms·
The problem is that even if the allocation cost (slab allocator) is zero, and even if the GC cost is zero, a high allocation rate on modern hardware effectively
by cpurdy 4y ago
The problem is that even if the allocation cost (slab allocator) is zero, and even if the GC cost is zero, a high allocation rate on modern hardware effectively flushes the cache lines at the rate of allocation, effectively reducing your L1 + L2 + L3 cache to 0MB total if your allocation rate is high enough. A slab allocator will be almost guaranteed to be allocating from a non-cached line, and thus the object init will evict a hot cache line every time.
But on top of that, neither the slab allocator nor the GC are free. They're fast, and they're very good, but they also make heavy use of the memory bus, thus competing heavily with the cache-thrashing already being caused by the objects being allocated.
This is why in Java, on a heavily threaded and heavily loaded process, you can see that the CPU cores are far from 100% utilization, but yet there's no blocked threads (and more threads than cores). In other words, once the memory bus saturates, the effective throughput of the CPU drops off and the system cannot make full use of the processing power available.
The relative cost of memory access has gone up 2 orders of magnitude in the past 25 years, and that's before the bus becomes saturated. When the bus is saturated, it can go up dramatically from there (see: queue theory).
- pveentjer 4y agoIf a Java thread is on the CPU, then the CPU is utilized, no matter if the CPU is just waiting for memory. That is the problem with utilization; you have no clue how busy the CPU actually is. IPC is more useful in these situations.