7 ms·
Faster Mac dev tools with custom allocators
- jeffbee 5y agoFor even more bang, rebuild LLVM with your custom allocator, profile guidance, and link-time optimization. If you're planning to invoke the compiler a million times, you might as well get it peak optimized.
- pdimitar 5y agoI'd love to read a detailed walkthrough about this!
- vlovich123 5y agohttps://llvm.org/docs/BuildingADistribution.html https://llvm.org/docs/BuildingADistribution.html
- liuliu 5y agoYeah, building Swift is not straight-forward. There are several repos needs to be checked out and coordinated in lock-steps. I usually starts with "swift/release/xxx" branches in llvm / cmark / swift, and call this to find what I am missing: https://github.com/apple/swift/blob/main/utils/build-script https://github.com/apple/swift/blob/main/utils/build-script
- bsaul 5y agoAre there any official benchmarks for swift compilation times ? I’m curious to see if they’ve been going up or down with the latest releases.
- pcwalton 5y agoPathfinder got enormous benefits (3x performance difference as I recall?) switching from Apple's default allocator to jemalloc. Apple's allocator has not kept pace with the competition, especially for multithreaded workloads.
- lukeh 5y agoWhat’s stopping Apple switching its default allocator?
- dwaite 5y agoI can't speak to Apple's reasons specifically, but usually it comes down to: 1. Degenerate cases with new allocator designs, e.g. being better for some workloads and not others 2. Bug-for-bug compatibility - applications which break due to dependencies on undocumented behavior or the old allocator memory structures. 3. Boundary conflicts - systems where an allocator change would mean allocation and free are hitting different implementations across module boundaries, as one module allocates memory for another to consume. Some systems and programming languages are more vulnerable to this sort of issue.
- vlovich123 5y agoA fourth one would be debugging support. They have a bunch of stuff in there (perhaps dated these days) to auto scribble on malloc/free, allocate guard pages etc. They probably could/should look into integrating a more modern allocator (mimalloc, jemalloc) and falling back to the old one when those features are needed (assuming they can’t bring them forward).
- pcwalton 5y agoMallocScribble is available as "opt.junk" in jemalloc [1]. As for guard pages, tcmalloc has TCMALLOC_PAGE_FENCE [2], and there is an issue [3] in jemalloc. In any case, Apple and others have invested hugely in LLVM AddressSanitizer, so the Electric Fence-like malloc debugging features are considered more of a last resort these days. [1]: http://jemalloc.net/jemalloc.3.html http://jemalloc.net/jemalloc.3.html [2]: https://chromium.googlesource.com/external/gperftools/+/gperftools-2.2.1/src/debugallocation.cc https://chromium.googlesource.com/external/gperftools/+/gper... [3]: https://github.com/jemalloc/jemalloc/issues/1664 https://github.com/jemalloc/jemalloc/issues/1664
- alberth 5y agomimalloc. Has anyone given Microsoft allocator (MIT license) a try? It appears to benchmark better than even jemalloc and others. https://github.com/microsoft/mimalloc#Performance https://github.com/microsoft/mimalloc#Performance
- meisel 5y agomimalloc is briefly mentioned in my article, and I was surprised that it worked on the mac for swapping itself in. I found the performance of it to be roughly the same as jemalloc when I measured
- OnlyMortal 5y agoWe use a modified (by the author) jemalloc in a heavily threaded server that runs on Centos. The reason was due to memory fragmentation over time and jemalloc reduced this to a level that we could live with. We’ve tried other allocators over the years but still jemalloc, or our version of to be precise, is the winner in memory usage and fragmentation. Edit: performance wise, gains can be had by doing your own memory pools on a per-thread basis and making use of stack objects rather than heap objects. Larger allocations for your own object pools and managing those reduces the calls to malloc also helps reduce memory fragmentation and your vmsize diverging from your rss.
- OnlyMortal 5y agoJust to add… std::vector can lead to fragmentation in heavily threaded applications. We found std::deque solved that issue.
- m_eiman 5y agoPerhaps there's a sandbox escape hiding in the workaround they did to get swiftc to use their jemalloc build?
- saagarjha 5y agoIt’s a build variable in Xcode telling it which compiler to use. At that point you have arbitrary code execution anyways as you control the build process.
- Jarred 5y agoMimalloc improved Bun’s performance on macOS by 10%.
- coldcode 5y agoCustom allocators have always been available for specialized needs, even the JDK ships with a number of options which each can be tuned further. I wrote a fast and safe memory allocator for the old MacOS that was briefly popular before MacOSX appeared which obsoleted it. But having built such a beast (and written the tons of test apps you need to ensure it works under all conditions) there is always room to optimize for needs that you can't employ in a generalized allocator. Like everything, you can't optimize for all cases and still be good enough for the average case.
- meisel 5y agoI'd be really interested to see benchmarks of Apple's allocator versus others, both in memory consumption and performance. Sometimes, Apple's version of things is worse in almost all use cases, but I'll reserve judgment and wait to see numbers.