4 ms·
I agree that there could have been a more satisfying conclusion, but it is worth noting that jemalloc isn't a panacea. I've seen issues similar to Svix in both
by mcronce 3y ago
I agree that there could have been a more satisfying conclusion, but it is worth noting that jemalloc isn't a panacea. I've seen issues similar to Svix in both Rust and C++ applications that were heavy on ephemeral allocations, and have fixed it by doing all of the following, depending on the specific process:
* Switching from libc malloc to jemalloc
* Switching from libc malloc to tcmalloc (dating myself a little bit)
* Switching from libc malloc to mimalloc
* Switching from jemalloc to mimalloc
* Switching from jemalloc to libc malloc
* Switching from mimalloc to jemalloc
Possibly others; I only want to list cases I'm 100% certain of.
Heap fragmentation is just a reality of some allocation patterns without a GC runtime.
One certainly can (and, in some cases, should) make their application more allocator-friendly, but - aside from some often-low-hanging fruit - this is a time-intensive process involving a bit of, for lack of a better word, arcane knowledge (I should inline all my fields and allocate on the stack as much as possible, right? Yes, well, except ...)
If you already have a halfway decent benchmark suite or workload generator, which you'll want for other purposes anyway, it's often a lot quicker to just try a few other allocators and select the one that handles your workload best.
- scottlamb 3y ago> * Switching from libc malloc to tcmalloc (dating myself a little bit) If you think of tcmalloc as an old crusty allocator, you've probably only seen the gperftools version of it. This is the version Google now uses internally: https://github.com/google/tcmalloc https://github.com/google/tcmalloc It's worth a fresh look. In particular, it supports per-CPU caches as an alternative to per-thread caches. Those are fantastic if you have a lot more threads than CPUs. I haven't checked if it's been adapted for the latest upstream kernel API, but there's also the idea of "vcpu"-based caches: basically rather than a physical cpu id, it's an (optionally per-numa-node-based) dense id assigned to active threads, so that it still works well if you have a small cpu allocation for this process on a many-core machine.
- jeffbee 3y agoI think we can safely say that jemalloc has much better marketing than tcmalloc. If you search on github there is only 1 real project with a bazel WORKSPACE that mentions tcmalloc ... the rest are either forks of tcmalloc itself or trivial toy programs that I personally put on github to share with someone. There are zillions of users of jemalloc and everyone has heard of it.
- scottlamb 3y agoAgreed. I think it's a wider phenomenon: Google is culturally uninterested in / incapable of marketing an open source project that's less ambitious / strategic than e.g. TensorFlow, Kubernetes, Go, Dart/Flutter, or (looking back a few years...) Angular. The next tier I guess is Abseil or Bazel, which have a nice website, nice docs, and some papers/talks but I think still aren't super widely used outside Google. I can't think of anything smaller than those that's gotten any marketing at all. Can you? tcmalloc at least gets internal changes regularly synced to github (since just a couple years ago iirc), vs. the ancient gperftools snapshot that's more widely known. And there are many other projects of potential interest to the outside world (Fibers...) that have been mentioned publicly but not open sourced at all.
- vlovich123 3y agoWe use it here at Cloudflare on every single machine as part of Workers. So that’s two major hyperscalers running large RAM multi tenant workloads. Jemalloc may more recognition in the broader community, but the largest workloads seem to be running mimalloc / tcmalloc (I don’t know what Facebook uses internally). The libc malloc probably has even more users than either as it’s the default allocator for iOS and Android.
- gmokki 3y agoAll of the allocators are improved over time (against their own their own goals). It is always important to specify the libc version and version of alternative allocators. Otherwise the results might not be reproducible a year later.
- pkolaczk 3y ago> Heap fragmentation is just a reality of some allocation patterns without a GC runtime. It is also a reality with a GC runtime, even a compacting one. Tracing GCs have acceptable performance only if the available memory is many times larger than the memory used. Technically this maybe isn't called fragmentation, but the effects are just as bad - the application uses much more memory than really needed.