20 ms·
I Improved My Rust Compile Times
- eikenberry 3y agohttps://web.archive.org/web/20240320042828/https://benw.is/posts/how-i-improved-my-rust-compile-times-by-seventy-five-percent https://web.archive.org/web/20240320042828/https://benw.is/p...
- tempaccount420 3y agoVery cool, still kind of disappointing the author was only able to get incremental compilation down to 4.5s, no matter the baseline. Hopefully Cranelift lands in stable at some point.
- dclowd9901 3y agoThat’s pretty good. Our frontend app built with TS with SWC recompiles in about that time.
- boxed 3y agoDepends on the code size though. Is this 1k lines? 10k lines? 100k lines? Just looking at the blog this seems like a fully static site that should be able to be built in milliseconds. Some input from the developer here on what takes all this time would be interesting.
- denysvitali 3y agoTheir benchmarks are based on a machine with 72GB of RAM and another with 128GB of RAM - not that it does matter, but I'm really curious if they're outliers or if I should bump my computer specs. That's a lot of Electron web apps running in parallel!
- mixedCase 3y agoSome people need it, some don't. When I upgraded to a 7950X I was unable to find an excuse to buy more than 32GB. "That one time a script compiled clang by default with LTO and OOMed" was not a good enough reason. If you use a lot of VMs, it comes in handy. If you do a lot of heavyweight compilation in parallel, same thing. But at least on Linux, popping 3 VS Code projects and 4 more browsers doesn't even hit 50% utilization.
- rakejake 3y agoFunnily enough, RAM I find I can use no matter how much I have, especially now that you have LLMs. I can run a quantised Mixtral 8x7B on 64GB RAM. But the 7950X is grossly underutilised, I don't know what to do with it.
- lolinder 3y agoShouldn't the quantized LLMs be using the CPU to good effect? I'd imagine your Mixtral is substantially faster than it would be on a weaker CPU.
- rakejake 3y agoI have a 4080 with 16GB of VRAM. I experimented with llama.cpp by offloading layers onto GPU and doing the remaining on CPU. I found that it gives me max tokens/sec if I set it to 8 CPU cores as opposed to the 16 available on the 7950X. I guess beyond that, the bookkeeping between the cores might be taking up more time than it is worth.
- mischief6 3y agothe answer i found is run gentoo and compile yocto projects.
- db48x 3y agoYea, very workload dependent. Reposurgeon needed hundreds of gigabytes of ram to convert the GCC repository from SVN to Git, which was a lot of fun, but most of the time I would be fine with 32GB.
- pcthrowaway 3y agoMy understanding is that processor performance is more important for rust compilation than available memory, but ultimately the atrocious compilation times are why I struggled so much with Rust.
- lolinder 3y agoDefinitely outliers. The Steam hardware survey [0] is obviously not a perfect measurement, but in its February results only ~3% had 64GB and only ~0.3% had more than that. I recently upgraded to 64GB to run quantized LLM models, but when I'm not using an LLM I rarely use more than 16GB, much less 32GB. [0] https://store.steampowered.com/hwsurvey https://store.steampowered.com/hwsurvey
- IshKebab 3y agoI have 32GB on my Linux laptop (plus 12GB swap) and regularly run out of RAM with just a handful of VSCode's open (it's generally the language servers using all the RAM) and a couple of dozen Firefox tabs (Firefox seems to be much worse than Chrome RAM-wise). If I run my actual work project, which needs a few gigs then it often just completely runs out of memory and hard reboots. Never have the same problem on Windows even though I only have 16GB there. Feels like Linux is just really bad at memory management. That may be skewing the Steam survey since almost everyone there will be using Windows. On Linux I wish I had 64GB.
- spookie 3y agoHave you tried zram for instance? There are a lot of things one can do to better utilise RAM on Linux. I believe Fedora comes ootb with a lot of things to make ram management quite good. Using another distro these days, but Fedora impressed me on that front. (edit: Try to use zram and zswap, check if your i/o scheduler is correct for your hardware, and increase your swapiness value. These should make your system behave better. For other performance improvements, use ananicy and irqbalance.)
- IshKebab 3y agoYes I have zram enabled. First thing I tried. Even if it is possible to fix this, it's still a horrible bug because you absolutely should not have to blindly tweak settings to stop a desktop computer from hard-rebooting when it runs out of RAM. I think this is a really fundamental problem on Linux though - it would require enormous changes and unprecedented cooperation between different groups for a proper fix. Much easier to say "you're holding it wrong".
- nindalf 3y agoI benchmarked Rusts performance on 2, 4, 8 and 16 cores on several widely used crates (https://arewefastyet.pages.dev/ https://arewefastyet.pages.dev/). Performance generally improved with the number of cores, but memory made no difference. All of these crates topped out at around 1GB and additional memory didn’t improve performance. So no, 72GB or 128GB makes no difference.
- drewtato 3y agoI agree that memory size isn't going to matter much, but memory speed should be noticeable, especially on large projects.
- deleted 3y ago[deleted]
- denysvitali 3y agoYes, maybe I didn't make it clear but I'm pretty sure this is not impacting for Rust. I was just wondering if it's normal to have these specs since my computer has probably "just" 32GB of RAM
- lifthrasiir 3y agoIf you are like me and opening too many browser tabs (and other applications, of course), you need a lot of RAM. My desktop has only 48 GB of RAM right now and a typical VM usage is around 80--100 GB. I will probably go for 128 GB or possibly 192 GB of RAM for the next desktop. Rust builds, in comparison, don't actually reach that far, mainly because it is proportional to the maximum number of parallel jobs. In my experience, 48 GB was more than plenty for 8 parallel jobs if I wasn't doing anything else.
- nextaccountic 3y ago> Note: Discord User @PaulH(and a couple others) let me know about this bug that prevents the server build from using the dependencies built for the previous compilation of the server/webassembly frontend build due to changes in the RUSTFLAGS. You can fix this by having cargo-leptos > 0.2.1 and adding this to you Cargo.toml under your Leptos settings block. I did not test with this enabled. > [package.metadata.leptos] > separate-front-target-dir = true This option is deprecated since cargo-leptos 0.2.3 (now it's enabled unconditionally), it's going to be removed in 0.3 https://github.com/leptos-rs/cargo-leptos/pull/216 https://github.com/leptos-rs/cargo-leptos/pull/216 https://github.com/leptos-rs/cargo-leptos/issues/217 https://github.com/leptos-rs/cargo-leptos/issues/217 https://github.com/leptos-rs/cargo-leptos/commit/b0c19a87cff2e0fb9bac0d9c6d274c38069ed2b5 https://github.com/leptos-rs/cargo-leptos/commit/b0c19a87cff...
- rui314 3y agoThe sold linker is no longer commercial software; it's been relicensed under the MIT license.
- mhaehnel 3y agoBut it is no longer updated, is it? Thanks anyway for the work on sold.
- School-Cotton 3y agoSold was just mold under a commercial license. Mold is still maintained and now licensed permissively.
- mhaehnel 3y agoThat's true but mold doesn't have MacOS support.
- rui314 3y ago> But it is no longer updated, is it? Essentially, yes, but that’s because Apple’s new linker is so much faster than it was. You can use it instead.
- nicce 3y agoOut of curiosity, do you know what technical choices Apple made to be so much faster?
- oefrha 3y agoIt’s parallelized. https://news.ycombinator.com/item?id=36218330 https://news.ycombinator.com/item?id=36218330
- rui314 3y agoI don't know the details, but once it is proven that achieving lld/mold/sold-level performance is technically feasible, it shouldn't be extremely hard to achieve similar success for macOS.
- dan_can_code 3y agoAs a front-end developer dabbling with rust, the compiler being slow is not a problem for me. Whilst I appreciate the benchmark, it's not likely I'll have 72GB of ram laying around to speed up the process. Yes, hot module reloading is nice and quick. But it reloads with errors included. The primary benefit of the rust compiler, at least from my point of view, is telling me what's wrong, where. It's a worthwhile sacrifice for a few seconds of my time, when at the end of addressing all of the obvious errors, I have something that works. I find this miles better than HMR.
- boxed 3y agoI use Elm at work, and the confidence from really strict compilation and typing is very worth it. With that said, Elm is super fast to compile too, which is very nice.
- pjmlp 3y agoI have a couple of projects where it is measured in minutes, on a dual core with 8GB and a SSD. If only devs with beefy machines are able to use Rust properly, that hinders adoption.
- guappa 3y agoDevelopers in USA: "nobody in the entire world has less than 32GB of RAM!"
- rtpg 3y agoLaptop + 32 gigs of RAM is a recipe for destruction of your wallet. At least for desktops you can kind of scope things out, but companies aren't really in the practice of distributing desktops anymore it seems.
- jakderrida 3y ago> Laptop + 32 gigs of RAM is a recipe for destruction of your wallet. While I'm not a developer of any sort, you couldn't be more on the nose here. I waited and waited for a Surface Laptop that has 32gb on ebay. When it came, price skyrocket to almost double the 16gb model. Twice more it happened and I literally gave up and bought the 16gb. I don't know what the retail premium from Microsoft is, but the reseller premium is truly insanity. Surely, Microsoft spent the time engineering it to fit 32gb and could reap more producing more of them.
- deleted 3y ago[deleted]
- winrid 3y agoThe coolest thing about this to me IMO is the comparison between the two CPUs. Is there a project that publishes benchmarks compiling projects across different CPUs? I know passmark is kinda the go-to but this would be cool. Someday I'll feel the need to upgrade from a 2700x...
- canu7 3y agoPhoronix does a lot of CPU benchmarks, including code compilation, but mostly focused on Linux. More result are also available in the OpenBenchmark page, which is also part of the same project. Take a look at the timed compilation test-suit: https://openbenchmarking.org/suite/pts/compilation https://openbenchmarking.org/suite/pts/compilation
- oynqr 3y agoHave you looked at https://openbenchmarking.org https://openbenchmarking.org?
- winrid 3y agoI haven't, thanks!
- beeb 3y agoSuper interesting! I would have loved to see included in the potential solutions the use of sccache too
- SushiHippie 3y agoLooking at some benchmarks online and the 7900x always beat the 5950x even though the 5950x has more cores. But in this article the 7900x performed worse than the 5950x, is it because of the core count, or are there some other factors?
- daghamm 3y agoFirst time I hear about mold. The performance looks good, but are there any plans for any other improvements such as LTO?
- rui314 3y agomold supports LTO out of the box.
- rob74 3y agoI really wish people wouldn't use AI-generated images for blog posts, it's so distracting! I clicked the link wanting to read the article, but instead spent several minutes looking at the image to find errors: what's shown on the screen (apparently it's "Jot Commeditin") looks like something between a hex dump, git blame, and occult incantations written in a long-forgotten alphabet. Probably you need a keyboard like the one on the desk (with keys arranged in a fractal pattern) to write something like that. And a mouse with a cable going nowhere...
- deanWombourne 3y agoI did the same! Then didn't read the article.
- doubloon 3y agoHe got three out of four of “manic pixie dream girl”
- deleted 3y ago[deleted]
- klabb3 3y agoSame, but what’s worse is that people never put in the caption that it’s AI nor what tool is used. Transparency goes a long way.
- KingOfCoders 3y agoFrom my experience, faster SSD have a suprising strong effect for some (simple) languages (probably not Rust), like Go [0] [0] https://www.octobench.com/ https://www.octobench.com/
- klabb3 3y agoIs it an unpopular opinion to say 75% of “a lot” is still “a lot”, plus now you have to keep track of the knobs and constantly monitor for and be conscious about not wrecking build times as you’re maintaining and developing. I’ve found in 10+ years of software development that speed of iteration cycle is highly correlated with productivity. Compile times is not the only input into this cycle time, but it’s a big one, and importantly, it’s within the control of the language tooling itself to solve. The human idle attention time of 1-2 seconds should be the gold standard to strive for, even if not always achievable. There seems to be quite a bit of cope around Rust build times in the community, which was natural a few years ago (a lot of people used to “blame” llvm, but it doesn’t seem to be as big of a culprit) but things are different now, no? Given the maturity, growing ecosystem and corporate investment, I would expect incremental build speedup to be prioritized, and steadily improving. But clearly it isn’t moving very fast in that direction. So why not?
- IshKebab 3y ago> I’ve found in 10+ years of software development that speed of iteration cycle is highly correlated with productivity. I agree, though in practice with many Rust programs that iteration cycle is actually not "edit, compile, run", it's "edit, save, wait for rust-analyzer to update" which is generally much quicker (even if it is sometimes still a little slow). In most cases, Rust's very strong type system means you spend a lot less time actually running your program. There are definitely exceptions though. E.g. if you're making a web app and editing CSS or whatever then the Rust compiler isn't going to tell you if you've got the colour wrong. > a lot of people used to “blame” llvm, but it doesn’t seem to be as big of a culprit It still is; there was a post here really recently where someone broke down the time spent in various phases. LLVM is still the vast majority, though apparently that is partly due to Rust generating a lot of work for it. > I would expect incremental build speedup to be prioritized, and steadily improving. It is. > But clearly it isn’t moving very fast in that direction. So why not? Because Rust's unit of compilation (a crate) is very large. In C/C++ it's a single file.
- rockwotj 3y ago> I agree, though in practice with many Rust programs that iteration cycle is actually not "edit, compile, run", it's "edit, save, wait for rust-analyzer to update" which is generally much quicker (even if it is sometimes still a little slow). This is one area that I think Dart did a good job on. They support both a JIT (I am not sure if it’s tiered) and AOT modes of compiling, the JIT targeted for development and AOT for production. This is the ultimate extreme of optimizing for both use cases, obviously it’s a ton of work, but I would love more languages to be like this. (I believe the JIT mode works directly on the source so there isn't an intermediate compilation step like Java, but it’s been a few years since I’ve used Dart) [1]: https://dart.dev/overview#native-platform https://dart.dev/overview#native-platform
- dgellow 3y agoI’m not sure I understand why editing an html template requires to recompile the rust project. I implemented a static site generator in rust for my own projects, html templates are simple liquid template files, it’s instant to recompile them. Same for stylesheets. Is the author generating the html from rust directly? If yes that sounds like a painful strategy
- richrichie 3y agoThe older i get the more i am convinced that there-ain’t-no-free-lunch applies widely, beyond financial markets. Software development: the art of redistributing aggregate lifecycle pain; who bears what, when and how much.
- OJFord 3y ago> i am convinced that there-ain’t-no-free-lunch applies widely, beyond financial markets That's amusingly ironic, since it doesn't originate in financial markets anyway! https://en.wikipedia.org/wiki/No_free_lunch_theorem https://en.wikipedia.org/wiki/No_free_lunch_theorem
- deleted 3y ago[deleted]
- rbalint 3y ago75% is really good even if it requires changing the toolchain. Have you tried https://github.com/firebuild/firebuild https://github.com/firebuild/firebuild ? It is a caching accelerator and it can cache the linking and the buildscripts, too in Rust builds. It can make your builds more than 90% faster, especially with low cores counts. https://balintreczey.hu/blog/improve-build-time-of-rust-java-and-intel-fortran-projects-with-firebuilds-new-release/ https://balintreczey.hu/blog/improve-build-time-of-rust-java... The Mac port is experimental, though.