8 ms·
Why does musl make my Rust code so slow?
- pedrocr 6y agoSwapping out the allocator for jemalloc would be my first try. It's easy to do and often results in better performance. 30x requires some kind of pathological case though.
- mrits 6y agoIn a commercial product I worked on I went against the vendors advice to try out jemalloc. It took a 100GB memory hold (that took 48 hours to happen) to staying steady at around 2-4GB and only peaking at 100GB for a few seconds a day. Same exact code but just swapped out the jemalloc at the command line.
- iou 6y agoSwap out the allocator https://users.rust-lang.org/t/optimizing-rust-binaries-observation-of-musl-versus-glibc-and-jemalloc-versus-system-alloc/8499 https://users.rust-lang.org/t/optimizing-rust-binaries-obser...
- asveikau 6y agoWithout familiarity with rust, I wasn't sure what they meant by "system allocator". Apparently that means libc's malloc. (Or HeapAlloc on Windows) So I guess they statically link jemalloc but can optionally use libc malloc.
- oconnor663 6y agoIt's the other way around now, though the post that iou linked is from an earlier period when things worked differently. By default today, Rust programs use the same allocator that C programs use, which I think is provided by libc on Linux, and you have the option of using a custom allocator, like jemalloc. Historically however, all Rust programs used to use jemalloc by default.
- tyingq 6y agoI suppose there's no option to link statically against glibc because of the implications with LGPL (static linking triggers additional license clauses).
- richardwhiuk 6y agoI think it's mainly that glibc has poor support for statically linking.
- tyingq 6y agoAh, okay. Searched a bit, and it apparently requires you to find your own with dns client lookup library, avoid dlopen(), set GCONV_PATH, and so on.
- aidenn0 6y agoFor a long time glibc was maintained by someone who was of the opinion that if you wanted to statically link 3rd party libraries, then you shouldn't be allowed to build binaries.
- asveikau 6y agoIt's interesting because there is an opposite cult (I believe some prominent plan9 and go people are adherents) that believes dynamic linking should literally not be a thing. I even recall reading some essays that introduction of a dynamic linker was some kind of tragic downfall in Unix history. I think the truth is that both have costs and benefits. Dynamic linking is good for security patches, memory and disk usage. But it creates new opportunities for problems, for example ABI breakage becomes more significant of a problem and needs a lot of care to avoid. People distributing code, be it to end users or app stores or on servers with things like chroots, jails and containers, need to carry their dependencies anyway negating some of the benefits.
- 6y ago
- andygrove 6y agoThanks! That really does seem to be the issue and I wouldn't have known about this, had I not asked. I will try this out and will update the blog post in ~8 hours time.
- masklinn 6y agoIME allocations is one of the main things making rust programs slow without diving into the more arcane stuff. So looking into unnecessary allocations and/or the performances of the allocator would be one of the first things to do (right after checking if you're compiling with optimisations). Given your CPU graphs, and the large number of cores, I expect musl's allocator simply has very poor behaviour with respect to multithreading (e.g. limited or no threadlocal arenas, size-classing, etc…) leading to a lot of crosstalk, extreme contention on allocations, etc...
- zokier 6y agoTbh allocations (and related memory management things) are often the low hanging fruit in big picture optimization in many languages, incl c++.
- BubRoss 6y agoThis actually someone asking and not an investigation and explanation. There isn't even a lot of due diligence to figuring it out - no profiling or resource usage other than CPUs. Also it is musl combined with docker causing a 30x slowdown. If something is running 30x slower from linking in a different libc, I'm guessing it should not be that difficult to narrow down the cause at least a little bit.
- fluffything 6y agoYeah, you should be upvoted. The author put zero effort into figuring things out.
- vardump 6y agoYeah, upvoted. 30x slower on 48 core (?) system sounds suspiciously like excessive lock contention (or some other shared resource). Non-NUMA aware allocator (or other code) might also contribute to the issue. Should be fairly easy to investigate.
- Animats 6y agoHe has so many layers in there it's going to be tough to find the problem. "Ballista is an experimental distributed compute platform, powered by Apache Arrow, with support for Rust and JVM (Java, Kotlin, and Scala)." Plus he's got Docker, the Rust library, musl, and jemalloc sometimes. There's no application. All this is just infrastructure. Musl doesn't do much on its own. But it does do stdio buffering. Could it be that the buffering system is making too many I/O calls, like flushing on every write?
- vardump 6y agoExcessive I/O is indeed another common issue. The number of times I've seen a slow system and discovered excessive number of flushes, often in something like logging system to be the root cause... Or gazillion of unbuffered 1 byte writes or reads.
- MiroF 6y ago
- termie 6y agorun a perf trace on both and see what jumps out
- MiroF 6y agoIf he's benchmarking on docker, I'm not sure that perf works in docker.
- wyldfire 6y agoYou may not even need to go that deep. Just strace (follow forks) and look what commands get exec'd. "Why does musl make my Rust code so slow?" But he's measuring mostly the compiler performance in "cargo build". Is he writing the same amount of data to disk in the same experiments? Seems like there's a lot of opportunity for some shallow investigation to find out more.
- pjc50 6y ago.. where's the profile output?
- underdeserver 6y ago...Not the Intel guy, if anyone else had to pause for a second.
- andygrove 6y agoI get that a lot!
- biesnecker 6y agoYou still being alive probably helps for disambiguation. :-)
- hinkley 6y agoAre you familiar with SwiftOnSecurity on twitter? Do you have any hobbies that would be out of character for Intel's Andy Grove? I think the world has room for a ficttionalized Andy Grove talking about how to cook french pastries, train bonsai, intermittent fasting, or preparing for a marathon.
- platinumrad 6y agoOne SwiftOnSecurity is already too many.
- hinkley 6y agoIt's a free country. You're allowed to be wrong.
- curtis3389 6y agohttps://lmgtfy.com/?q=musl+libc+performance https://lmgtfy.com/?q=musl+libc+performance I'm fairly sure musl is used because it's a smaller alternative to glibc, so it's better suited to Docker containers. It's goals are security, not performance.
- jessermeyer 6y agoFor those curious, Musl's malloc implementation is currently being re-written for higher performance and robustness, see https://github.com/richfelker/mallocng-draft https://github.com/richfelker/mallocng-draft
- liuliu 6y agoDo you have any extra readings on the rationale of building their own malloc rather than integrating mimalloc or jemalloc?
- The_rationalist 6y agoThe NIH syndrome obviously, otherwise it would have been addressed at the start of the Readme
- harrygeez 6y agowell one of the stated goals of musl is to be simple and correct, and all those mallocs are anything but simple
- littlestymaar 6y agoThat's true for jemalloc, but mimalloc is pretty simple. The reference paper is pretty short and really accessible and IIRC the implementation is around 5klocs. I doubt musl's implementation would much simpler than this.
- dalias 6y ago5kloc is about 10x larger than musl's existing (old) malloc in source lines. I suspect lots of that is low code density, comments, etc. I have to lookup what exactly mimalloc is/does every time someone mentions it, because the readme/documentation isn't very descriptive except discussing extensions outside the normal API. I didn't have time to dig through this again today. But I did look at it in some depth on several occasions in the past and it really wasn't suitable for or comparable to what we're doing in musl.
- renewiltord 6y agoPost was not very illuminating. Very little content. It's pretty much a "if musl is slow, it may be the allocator (eom)" which fits in the headline and would have saved me the click.
- dalias 6y agoWe'd be happy to address specific problems on the mailing list. I believe it's a known issue that the Rust compiler is making really heavy use of rapid allocation/freeing cycles, and would benefit from linking a performance-oriented malloc replacement. Doing so is inherently a tradeoff between many factors including performance, memory overhead, safety against erroneous usage by programs, etc. One statement in your post, which some readers pointed out was apparently added later, "Others have suggested that the performance problems in musl go deeper than that and that there are fundamental issues with threading in musl, potentially making it unsuitable for my use case," seems wrong unless they just meant that the malloc implementation is not thread-caching/thread-local-arena-based. The threads implementation in musl is the only one I'm aware of that doesn't still have significant bugs in some of the synchronization primitives or in cancellation. It's missing a few optional and somewhat obscure features like priority-ceiling mutexes, and Linux doesn't even admit a fully correct implementation in some regards like interaction of thread priorities with some synchronization primitives, but all the basic functionality is there and was written with extreme attention to correctness, and musl aims to be a very good choice in situations where this matters.
- georgianar 6y agoEmail this great herbal doctor who cured me from herpes virus and also brought my lover back via his email robinsonbucler@gmail. com