4 ms·
FWIW, that synthetic benchmark was reflective of some real world code we were deploying. Using malloc/free for one function led to something like a 2x performan
by celrod 3y ago
FWIW, that synthetic benchmark was reflective of some real world code we were deploying.
Using malloc/free for one function led to something like a 2x performance improvement of the whole program.
I think it's important to differentiate between malloc implementations/algorithms, just like it's important to differentiate between GCs.
E.g., mimalloc "shards" size classes into pages, with separate free lists per page. This way, subsequent allocations are all from the same page. Freeing does not free eagerly; only if the entire page is freed, or if we hit a new allocation and the page is empty, then it can hit a periodic slow path to do deferred work.
https://www.microsoft.com/en-us/research/uploads/prod/2019/06/mimalloc-tr-v1.pdf https://www.microsoft.com/en-us/research/uploads/prod/2019/0...
Good malloc implementations can also employ techniques to avoid fragmentation.
It's unfortunate that the defaults are bad.
But I confess, compacting GCs and profiling the effects of heap fragmentation (especially over time in long running programs) are both things I lack experience in. Microbenchmarks are unlikely to capture that accurately.
- yxhuvud 3y agoIf we are talking long running multithreaded processes, then libc malloc have issues that isn't even fragmentation. There are many workloads when it seems to forget to return whole empty pages and just accumulate allocated areas despite them being totally empty.