13 ms·
Speeding Up the Rust Compiler
- eb0la 7y agoLooks like if you're using an 'old' release (<=1.36 , like me) it is time to upgrade. Performance improvements in this post are from november'19 to december'19 - IMHO the november version _already_ had impressive optimizations.
- ChrisSD 7y agoNote that `cargo check` is faster than doing a full compile. Also I use the `rust-analyzer` language server for IDE integration to catch errors as I write them. Between the two, my workflow usually avoids the need for actually compiling a binary until I'm ready to run tests.
- EugeneOZ 7y ago'cargo clippy' is another good option (also faster than full compilation).
- simias 7y agoWhile developing I completely agree with you. However when I'm debugging it tends to get messier in my experience, I often want to make small changes and see how they influence the symptoms. In this situation you have to do full builds every time, and if your application is performance-critical enough you may not have the option of using the faster non-optimized builds. Admittedly part of the blame falls on me, I'm a big fan of printf-debugging and I tend to only use debuggers as a last resort.
- est31 7y agoAlso, due to this bug, adding a single printf line causes everything after it in the file to be recompiled without there being a need: https://github.com/rust-lang/rust/issues/47389 https://github.com/rust-lang/rust/issues/47389
- ChrisSD 7y agoAh, I'm quite old fashioned when it comes to debugging, so I'll attach a debugger and maybe even read through the generated assembly if need be. I've also been trying to make more use of static asserts where possible. If I can ensure both me and the compiler have the same understanding of the code then there's hopefully less space for errors at runtime. But of course this is never going to eliminate the need for debugging, even if (and this is a big IF) the compiler knows my exact intent (i.e. my reasoning might be wrong somewhere, perhaps very subtly so). And static asserts are still quite hacky. All that said, I totally agree that there are times when faster compiles are really useful.
- banachtarski 7y agoYou’ve just given the reason I often use to explain why printf debugging is something I’m not a fan of. For native development, printf debugging takes a backseat to proper instrumentation and debuggers just for sheer productivity reasons IMO. The main thing it gives is a serializable log for debugging multithreaded bugs. However, for just viewing state during a run, I think people should just learn to use their debugger (which can inject prints and watches on the fly, no need to recompile).
- simias 7y agoI agree with you, the main reason I stick to printf debugging is because I can't be bothered to change my habits, it's not a very sensible choice. That being said, if I had to defend myself, I'd point out that printf debugging has a few advantages: - It's often more lightweight and less intrusive than debuggers, which makes it less likely to encounter an "heisengbug" that disappears when you attempt to debug it. This is especially true for timing-sensitive bugs (which occur in multithreaded code as you point out, but not only). - Debuggers are often environment-specific, if you change language, environment or even simply editor you might have to re-learn how to use your debugger. Printf will always be here for you. - I do a lot of work in embedded environments, including low level and bare metal stuff (bootloaders, drivers etc...). While these days there's generally some debugger support available for these targets it's often more limited or much more intrusive. If I put a breakpoint in an interrupt handler I basically freeze the entire kernel when it triggers which can sometimes do more harm than good if I'm trying to figure out what's going on. And again, in these environments the debugging solutions are often proprietary and sometimes quite expensive.
- dakom 7y agoDoes the VSCode integration support this? Though for webassembly I need the generated wasm to exist and be loaded before I can really see what's happening :\
- steveklabnik 7y agorust-analyzer includes a VS: Code plugin.
- bluejekyll 7y agorustc might never be as fast as the Go compiler because the language has so many additional features, but it always makes me so excited to see the continued progress in making the compiler ever faster. Thank you for the hard work here! Btw, mentioned in the article are these tools: “ All of the above improvements (and most of the ones in my previous posts) I found by profiling with Cachegrind, Callgrind, DHAT, and counts, but there are plenty of other profilers out there.” Does anyone have any good resources on using these with Rust, or just in general with C or C++?
- LordHeini 7y agoI think rustc will never be faster than go or even java. Go has made quite a few language concessions to be fast to compile and is designed with that in mind. And it really shows, using auto reloading in Go feels like using Ruby or Python which is just great. From my usage standpoint Go has basically no compile time at all. And awful compile times can be a huge hindrance to productivity. I used to do web dev in Scala, but waiting for the sleepy compiler is one of the reasons is switched to Go. Scala is a nice language but the long compile times in combination with the vm/jetty cycle feels quite a bit slower than Rust.
- pjmlp 7y agoDelphi, F#, ML, Ada, D prove otherwise. What they have going for them is not depending on LLVM.
- nicoburns 7y agoIndeed. LLVM seems like a big blocker for fast compiles, although I suspect Rust may additionally need some higher-level optimisation passes. The alternative cranelift backend here has been making slow but steady progress, and seems to be approaching a useable state: https://github.com/bjorn3/rustc_codegen_cranelift/issues https://github.com/bjorn3/rustc_codegen_cranelift/issues
- thegeekpirate 7y ago
- frenchman99 7y agoOne of the things with Rust is that while the compiler could be considered slow, once your Rust code compiles, if you stay away from `unsafe` code and `unwrap()`, the code is usually bug free (apart from logic bugs that no compiler could catch). At least that's been my experience with Rust.
- abjKT26nO8 7y agoAlthough Rust's type system does catch a lot, I wouldn't say it's that effective. I, for one, will usually make an off-by-one error or forget that I left a stub somewhere and didn't come back to write a proper implementation. But it's easily caught by the most rudimentary tests; so you don't have to bother yourself with writing them as elaborate as people usually do with e.g. Python.
- frenchman99 7y agoAn off-by-one error is a logic error.
- laumars 7y agoThat doesn't mean it isn't _also_ a software bug given that the result is still unintended / unexpected.
- bluejekyll 7y agoA lot of those basic issues, clippy often catches. One thing I enabled recently to make sure it’s caught before release, is checking for spurious uses of dbg! and unimplememted!, as well as println! in production code. For off by 1 errors, in Rust it’s often better to turn to iterators where appropriate than to say using indexes (it’s also more efficient in most cases). Clippy can also catch those issues.
- abjKT26nO8 7y agoIn my experience, all Clippy does is find style issues. Perhaps it's more effective in other projects, but in the code I wrote all by myself it didn't find a single error (as in bug). That's not to say Clippy isn't useful. I value style consistency (in this case consistency among the general pool of all Rust programmers and not inside a single project). Of course, YMMV. > For off by 1 errors, in Rust it’s often better to turn to iterators where appropriate than to say using indexes (it’s also more efficient in most cases). The off-by-one errors I made weren't related to indexing collections. I don't remember anymore what it was exactly, it got caught by the very first tests I wrote before even trying to use anything. I do use iterators whenever they make sense and I write new iterators whenever it makes sense. Still, good note from you.
- xiphias2 7y agoOne thing that was not addressed is why the effort is not put into making the parallel compiler default. It would give a 8x speedup on a developer machine compared to 10% speedups from these optimizations, especially when AMD releases the Zen 2 architecture for laptops.
- swsieber 7y agoAny gains in the single threaded model (usually) carry over to the parallel model. And maybe there is a separate effort on making the parallel compiler working better. Maybe they are working on it in parallel.
- xiphias2 7y agoI understand that there's a separate effort, but I don't see many blog posts about it, and I see that as something that should be prioritised. Rust was partly created because most of the transistors on the computers are heavily underutilized, and multiprocessing with C++ correctly is extremely hard. Rust compiler is written in Rust, so it would be a perfect showcase of taking advantage of the multi-processing safety of the language. I know that there are global variables in the compiler that the compiler team is getting rid of, but at this point that should be the main focus, as I see that most of the easy huge gains of single-threaded improvements are over.
- ronlobo 7y agoHaha, nice pun there!
- CJefferson 7y agoThe problem is lots of bits are hard to parallelise. Also, rust already does quite a bit in parallel, it can usually fill my 6 CPUs when building multiple crates.
- xiphias2 7y agoI see, I'm looking at the Rustc parallel meeting videos right now, I just wish there was a more organized blog for that team, or at least it would be easier to find the meeting notes. https://www.youtube.com/watch?v=Wh20eXfMOSk&t=8s https://www.youtube.com/watch?v=Wh20eXfMOSk&t=8s I found the meeting notes: https://hackmd.io/_1S8_ChMSa2N8mRw6EsGPA https://hackmd.io/_1S8_ChMSa2N8mRw6EsGPA
- fnord77 7y agoWe have a fairly small but complex library written in Rust. a debug compile takes about 5 minutes from clean. Release takes 18 minutes. (1.38.0)
- ComputerGuru 7y agoHow much of that is linking? I have simple project where the linker used to take over an hour, but ld.lld is much faster.
- gameswithgo 7y agoDo you happen to be using a lot of trait bounds? I had a small library that was slow due to this, and was able to refactor from ~3 minute builds to ~6 seconds by grouping the trait bounds into marker traits. There is a compiler flag you can use to list where time is being spent and it showed all my time was spent dealing with those trait bounds. There is an open issue to make that not necessary. If your time is spent linking, you can swap the lld linker in
- kibwen 7y agoThough note that LLD isn't supported on Mac, so no luck there.
- pcwalton 7y agoThat seems like a bug. Please file it.
- kzrdude 7y agoIs the library still small if we count its dependencies too?
- derefr 7y ago> The PR gave some very small (< 1%) speed-ups on the standard benchmarks but sped up a microbenchmark that exhibited the problem by over 1000x, and made it practical for procedural benchmarks to use tokens. Does this imply that existing benchmarks weren’t catching the problem here because they were avoiding making use of a feature because it was too slow? That seems like a strange way to write a benchmark, especially if tokens were actually in common use in the compiler itself.
- pdpi 7y agoGiven that we're talking about token concatenation in macros, I doubt that's going to show up in that much day-to-day code, so the small gains come from benchmarks using stdlib functionality that uses this feature. The bit about benchmarks using tokens sounds to me like he's talking about harness code, rather than the test subject code. You can speed up running the benchmarks without necessarily affecting the benchmark results themselves.
- naniwaduni 7y agoIf your benchmark harness has performance properties not reflected in your benchmark suite, your benchmark suite is incomplete. Benchmark harnesses are real programs too.