19 ms·
A possible new back end for Rust
- pizlonator 6y agoThis is really great. The world needs more diverse compiler tech. The llvm monoculture is constraining what kind of compiler research folks do to just the things that are practical to do in llvm. I particularly suspect that if something like Cranelift gets evolved more then it will eventually reach throughput parity with llvm, likely without actually implementing all of the optimizations that llvm has. It shouldn’t be assumed that just because llvm has an optimization that this optimization is profitable anywhere but llvm or at all. Final thought, someone should try this with B3. https://webkit.org/docs/b3/ https://webkit.org/docs/b3/
- pjmlp 6y agoThis is also why I think it was great that Maxime eventually graduated into GraalVM. Another tool for compiler research using modern approaches with type safe languages.
- The_rationalist 6y agoWho is Maxime?
- fluffything 6y agoIsn't GraalVM completely tied to LLVM bitcode, and therefore has all the same problems that LLVM has ?
- anp 6y agoIsn’t that just the (née sulong) llvm frontend? IIUC GraalVM is deeply dependent on OpenJDK internals.
- iamrecursion 6y agoNot at all. There’s an LLVM bitcode interpreter built on top of GraalVM, but the VM itself is heavily reliant on the internals of OpenJDK.
- pizlonator 6y agoNot even remotely.
- edwintorok 6y agoWorth mentioning other alternative small backends: http://c9x.me/compile/ http://c9x.me/compile/
- ddevault 6y agoI use qbe, it's great. Here's a mostly feature-complete C11 compiler based on qbe: https://git.sr.ht/~mcf/cproc https://git.sr.ht/~mcf/cproc
- hawski 6y agoI see that cproc is under quite heavy development, but qbe had last commit at the end of November. Is it considered feature complete? I heard about it some months ago and was quite interested in QBE, but it did not enjoy high tempo of changes. It may be considered advantage, I know too little to judge.
- ddevault 6y agoIt's complete enough to compile C11 programs - to me, that's as good of a benchmark as anything. The main thing qbe is missing for cproc's purposes is inline assembly and VLAs. DWARF support would also be nice, but no one seems to care enough to do the work yet.
- bluejekyll 6y agoWhile the GP doesn’t state this as an advantage, the Rust community would benefit from a fully Rust toolchain.
- VHRanger 6y agoNote that the LLVM monoculture came about because of how much of a pain GCC is to work with. And GCC being a pain to work with is a deliberate decision by Stallman to avoid his baby being expanded upon by corporations
- ddavis 6y agoIt's unfair to rms to say that. He would be happy for corporations to use and contribute to any project associated with the GNU project (like GCC), if everyone wanted to play along in GPL land (which of course isn't reality).
- yjftsjthsd-h 6y agoI think it's partially fair; gcc, in order to make it impossible to add proprietary add-ons, deliberately has an non-modular architecture, which makes it hard even for open source extensions to exist.
- CJefferson 6y agoUnfortunately not. For years people wanted gcc to output a nice parse tree for C++, which would have been plenty useful for open source text editors, but was banned by RMS as it would also be useful for closed source systems.
- phkahler 6y agoNot sure, but I think his concern was more about the introduction of opaque steps being introduced in the compiler and becoming something people depend on. A weak analogy might be nVidia drivers on linux, imagine a new arch where part of the toolchain is a closed blob. It turns out that hasn't happened yet with LLVM and allowing such things under LGPL may have worked.
- craftinator 6y agoI agree with this point of view; the three E's, Embrace, Extend, Extinguish, are already rampant in the compiler industry, and I see his choices as sacrificing ease of use for more transparency.
- Rochus 6y agoNow it's called monoculture? Rather strange. But anyway: in your terms you're just replacing LLVM monoculture by Rust monoculture, isn't it?
- yjftsjthsd-h 6y agoYes, it's a monoculture when the majority of all compiler work/research is happening on one compiler chain. (I feel like GCC is still competitive enough to keep up some competition, but Clang does have a lot of backing.) And yes, if we made a rust replacement and that somehow eclipsed all other compiler suites it would be a monoculture and be bad, but that's unlikely and creating an alternative to the most popular option reduces monoculture issues by adding more options.
- mratsim 6y agoLLVM has a library and modular approach which makes it easier for people to contribute just in their area of expertise instead of having to find their way in the hundred of thousands of line of GCC.
- Rochus 6y agoSo if we join forces and create a reusable compiler backend so not every compiler writer has to implement the same optimizers and code generators over and over again, then this is bad because it's a monoculture? How strange is that? To me, it sounds more like political propaganda from a few idealists who want to justify why - instead of participating in a joint project - they want to develop everything themselves from scratch in their favourite technology. For this there is, nota bene, also a common term: "Not invented here" syndrome.
- yjftsjthsd-h 6y ago> So if we join forces and create a reusable compiler backend so not every compiler writer has to implement the same optimizers and code generators over and over again, then this is bad because it's a monoculture? How strange is that? Why is that strange? You now have a diverse set of frontends and a monoculture on the backend. A world with Chrome, Chromium, Edge, Brave, and the Yandex browser is still a browser engine monoculture.
- benibela 6y agoThere are non-llvm compilers. FreePascal for example has its own x86, arm, mips, sparc and powerpc backends
- swagonomixxx 6y agoDevil's advocate: more diverse compiler tech will mean a more fragmented community and a larger probability of divergence across implementations. People think the C compiler community is dominated by GCC and Clang, and it is, but there are literally 1000s of implementations out there in the wild. Most are necessary, because we need code generated for some obscure processor architecture that's completely proprietary, but you can create that "backend" in LLVM itself - it's a new target architecture instead of e.g x86. The great thing about LLVM is that it's effectively the quickest (and probably the best) way to generate machine code without putting in too much effort, for a language. Whether that language be a research language or an existing industry language (say, C), that kind of establishment is hugely valuable. A great example of a good monoculture is the Go monoculture. Sure, there's gccgo, but the proportion of people using that vs. the reference implementation is minimal, and that reduced fragmentation is actually a good thing for practitioners (which most engineers are, not PL researchers).
- dmix 6y agoWhat about the usecase provided in the blog post, where one is used for fast debugging/dev builds but you use LLVM for production releases? Basically: why not both? I'm guessing this could create some divergence in terms of what is supported by the compiler but I'm curious how much that would matter in reality - for day-to-day serious project development. I'm not familiar with language dev at the compiler level, so I'm curious to hear if that's practical or sane.
- pizlonator 6y agoLlvm is absolutely not the least effort for generating machine code. In many settings, it takes a fraction of the effort of integrating llvm to create a template compiler that goes straight to machine code. In many other cases, your best bet is to have your compiler emit C and then feed that to a C compiler of your choice. It’s good to have divergence. Competition is good. Otherwise people stop trying new things.
- bluGill 6y agoDepends on your goals. Writing a front end, optimizer and backend quickly gets to more work. I can write a c++ compiler in a few months. It won't be good and to make it good would be many many years of work. If I write a llvm backend it might take a little longer (I doubt it), but I automatically get all the optimizations llvm has plus a good front end that doesn't have bugs in obscure corner cases. (not claiming llvm is perfect but there will be less bugs)
- cfv 6y agoI still remember when Clang bringing LLVM along was seen as SO OUT THERE and I'm just mentioning it because I find it weird to be old enough to see fads in system languages come and start to go. Just curious, do you have any examples of this "limitations" you speak of? Sounds like a very interesting read.
- stjohnswarts 6y agoHN isn't the place to go for conservative opinions on compilers :)
- fnord123 6y agoLLVM's MCJIT library is 17MB. If you have a language that you want to JIT and you thought you could embed your language like lua (<100k), Python (used to be ~250k but now <3M), you're looking at almost 20MB out of the gates. Not ideal! Also if you want to use llvm as a backend for your project and expect to build llvm as part of a vendored package, the llvm libraries with debug symbols on my machine was about 3GB. Also not ideal.
- Someone 6y agoAs an example, WebKit had an LLVM-based JavaScript optimizer in 2014 (https://webkit.org/blog/3362/introducing-the-webkit-ftl-jit/ https://webkit.org/blog/3362/introducing-the-webkit-ftl-jit/), but dropped it for another one in 2016 (https://webkit.org/blog/5852/introducing-the-b3-jit-compiler/ https://webkit.org/blog/5852/introducing-the-b3-jit-compiler...) In broad strokes, LLVM chooses to optimize for generating good code for statically compiled code more than for, for example, memory usage, compilation speed, or ability to dynamically change compiled code. That doesn’t make it optimal for JavaScript, a language that’s highly dynamic and often is used in cases where compilation time can easily dwarf execution time.
- pizlonator 6y agoWorth noting that B3’s biggest win was higher peak throughput. It generated better code than llvm. It achieved that by having an IR that lets us be a lot more precise about things that were important to our front end compiler. It’s not even about what language you’re compiling. It’s about the IR that goes into llvm or whatever you would use instead of llvm. If that IR generally does C-like things and can only describe types and aliasing to the level of fidelity that C can (I.e. structured assembly with crude hacks that let you sometimes pretend that you have a super janky abstract machine), then llvm is great. Otherwise it’s a missed opportunity.
- enos_feedler 6y agoWe are also seeing MLIR emerging as a compiler framework and LLVM being a dialect of that. This is happening within the LLVM project itself. From this point, it may be easier to write compilers without bringing in all of LLVM with it.
- amelius 6y agoDoes it support JIT compilation, i.e. specialization at runtime?
- steveklabnik 6y agoCranelift has a JIT, but I am not sure what the status of it is as a rustc backend.
- Koshkin 6y agoI've been wondering lately if the modern compilers should all be using C as the intermediate language (or some language-specific code optimization opportunities could be lost if they do that).
- edwintorok 6y agoThe semantics of C aren't very well defined, there is a lot of ambiguity in the form of undefined and implementation defined behaviour. This ambiguity is often needed to build an efficient optimizing compiler. When you have a higher level language with more accurately defined semantics, running it all through C would risk introducing undefined behaviour. With an IR you can control and define the semantics more closely to what your language needs.
- ansible 6y agoI've got to wonder if any of the existing intermediate representations would be appropriate with other programming languages.
- steveklabnik 6y agoThis is true to varying degrees, you could say that LLVM-IR and Java bytecode are two examples of this in action.
- jfkebwjsbx 6y ago> When you have a higher level language with more accurately defined semantics, running it all through C would risk introducing undefined behaviour. No, it wouldn't. When you target C you need to write a proper backend for its abstract machine, rather than naively rewriting code, of course. The C abstract machine is a fine IR, specially the later editions of the standard.
- epage 6y agoWhile a C backend is great for compatibility, is it a sufficient IL to express everything? For example, Rust has some extra guarantees with aliasing that I'm unsure if C or C extensions support yet that could offer greater optimizations (currently not fully being used due to bugs in the LLVM backend).
- crad 6y agoWhile it appears that cg_clif is faster to compile, does it provide any performance benefit compared to cg_llvm? Are the compiled binaries as fast as llvm compiled binaries? If not is the use-case for development purposes only?
- __s 6y agoCorrect, cranelift is meant for faster development build cycles https://github.com/bytecodealliance/wasmtime/blob/da02c913cc5a71de955b071a05bc157de39b20be/cranelift/rustc.md https://github.com/bytecodealliance/wasmtime/blob/da02c913cc...
- wscott 6y agoFrom the article, it is pretty clear that the resulting code is not as optimized as the LLVM backend. I didn't see any claims of how much slower it would be, but clearly that will vary greatly. Fast to compile is still really handy while developing.
- liquidify 6y ago>>"That’s Bjorn3, he decided to experiment in this area whilst on a summer vacation, and a year & half later single-handedly (bar a couple of PRs) achieved a working Cranelift frontend." Is this guy human? This is amazing, and this guy should be given an award.
- korpiq 6y agoThis feels welcome to me. I tend to think a language needs multiple independent implementations that only share the same source language spec, in order to really tear a clear spec apart from the quirks of any particular implementation. I find Rust (the spec, though also the implemenration) quite safe and practical (a balance). It deserves some independent implementations to secure a long and stable future. On the other hand, I want to use it on non-ARM embedded platforms, where current cross-compilation through C produces unusably big binaries. I dream this might increase hope for that, too, eventually.
- thesuperbigfrog 6y ago>> I find Rust (the spec, though also the implemenration) quite safe and practical (a balance). It deserves some independent implementations to secure a long and stable future. Where is the Rust spec? Unless something happened really quickly that I was not aware of there is only the implementation.
- steveklabnik 6y agohttps://doc.rust-lang.org/stable/reference/ https://doc.rust-lang.org/stable/reference/ is the closest thing we have. It is not yet complete.
- thesuperbigfrog 6y agoThank you! I look forward to the day when there is a spec, but I was surprised to see it mentioned and was wondering if I missed something big.
- jlebar 6y agoIf there are any rust people here, you've probably considered that you can speed up your debug llvm builds by enabling some optimizations. SimplifyCFG comes to mind, but, like, you can experiment. I presume the reason you haven't is because you want to preserve debug info, and llvm isn't great at that when optimizations are on.
- the8472 6y agoYou can customize the debug profile or create an intermediate profile between release and debug in your Cargo.toml. Debug info and optimization levels can be configured separately. If by speed up you mean compile times and not runtime behavior then there's also some unstable compiler flag that allows adding specific llvm passes.
- mttyng 6y agoThis is awesome. It doesn’t even seem that long ago when Boa was started! Man, time flies and people do great things. Kudos to the author and co-contributors for what Boa has become.
- gok 6y agoNovel compiler backends are a super cool idea, but I don't think it's going to help Rust compile speeds as much as this posts suggests. The complexity of Rust's type system puts a pretty high lower bound on compile times because of work the front end needs to do. Plain C compiles quickly even with an LLVM backend, for example.
- steveklabnik 6y agoWhile the type system does add to compile times, profiling generally doesn't show that it's the current limiting factor for compile times. Additionally, tools like rust-analyzer will give you type errors pretty much instantaneously, though of course that work is not finished. Also of note, this blog post isn't speculation; they posted numbers from actually doing it.
- gok 6y agoA 30% speedup is nothing to sneeze at, but it's not putting Rust within spitting distance of Go or C for similar amounts of code.
- steveklabnik 6y agoAbsolutely. I don't see where anyone is claiming that.
- dtolnay 6y agoRustc normally spends way more time in LLVM than in the frontend. Rust parsing and type checking are very fast in comparison to LLVM's codegen. Here is a chart from last September showing where the time goes in compiling a large Rust codebase (rustc itself): https://gistpreview.github.io/?74d799739504232991c49607d5ce748a https://gistpreview.github.io/?74d799739504232991c49607d5ce7... (Scroll down to the large horizontal bars once dependencies have been built.) (Sorry if GitHub is down at the moment; try later if it doesn't load.) The blue part of each bar is time in the frontend, the purple part is time in LLVM. The largest bar (rustc) spans 105 seconds in LLVM out of 140 total, or 75% in LLVM. Many of the subcrates are even more dominated by LLVM time, for example look at rustc_metadata or rustc_traits where >95% of compile time is spent in LLVM.
- tyrion 6y agoThanks for nice article! Hoping the author reads the comments, I would like to leave an, hopefully useful, feedback. It would greatly improve the reading experience of your blog if you could make clickable the footnotes/references. For example when you say: > I’ve taken the chart from the 2016 MIR blog post[3] I have to scroll to the end of the page to find the blog post (and then scroll back to resume reading). If [3] were clickable it would be great. It would be even better if [MIR blog post] were an actual link itself.
- The_rationalist 6y agoIf I remember correctly, mozilla had layoff a few months ago and the developper(s) of cranelift were in the bag. So is anybody currently paid to develop this backend? Without human resources I fail to see how this would keep up with truly supporting rust. As an aside, while the goal of faster build time is an important one, for completeness sake, I must tell that the mentality of rustc developers to be backend agnostic (an ideal) come at the cost of preventing rustc from adopting most llvm attributes and this fact is at the advantage of c++.
- Waterluvian 6y agoWhen writing a language like Rust, is the biggest challenge simply deciding what Rust's features and behaviors should be? And implementing the syntax and Rust -> LLVM compiler is really just a chore for the individuals who are super familiar with the implementation of these languages? Or is the technical implementation also genuinely challenging and non-obvious?
- kevinmgranger 6y agoThe concept of lifetime management is relatively novel and uncharted territory, if I understand correctly. There's only some prior art. So implementing that must have been an adventure and a half. And while I'm sure the folks who work on these languages are wonderfully intelligent people, let's dispel this notion that you need to be a super genius to implement a compiler or something like that! It seems magical, like one of the hardest things you could program-- but take a look through crafting interpreters, if you will: http://craftinginterpreters.com/ http://craftinginterpreters.com/ "Nothing is particularly hard if you break it down into small jobs." - Henry Ford
- Waterluvian 6y agoI walked through the Java portion of Crafting Interpreters and indeed, they can be much simpler than you might imagine in your early years. I was more just paying a compliment than suggesting that maintainers are superhuman. Everyone who ships productive code is a wizard and you can be too! I modified my original question to avoid a potential distraction from what I want to talk about. Thanks!
- xscott 6y agoDo you have any links for the prior art? I'm sincerely interested, thank you if you do. I'd also be interested in any articles (or even blog posts) that describe Rust's process for that from a compiler writer's point of view.
- tadfisher 6y agoA language like Rust aims for "zero-cost abstractions", which means the features and behaviors of the language must be evaluated in the context of the implementation.
- andrewprock 6y agoThe thing that struck me most about the article was this quote from the Rust Survey (2019): “Compiling development builds at least as fast as Go would be table stakes for us to consider Rust“ Go was designed from the ground up to have super fast compile times. In fact, there are some significant language issues related to that design decision. Using one of the primary design goals that impacted language structure as "table stakes" is almost certainly going require a lot of effort with some serious unintended consequences. Improving compilation times sounds good. Aiming high is good. But reaching "best of breed performance" is major initiative.
- pjmlp 6y agoIf you mean generics, D, Delphi, Ada and plenty of other languages prove you can have them and still be pretty fast.
- andrewprock 6y agoI mean interface{} https://golang.org/doc/effective_go.html#interfaces_and_types https://golang.org/doc/effective_go.html#interfaces_and_type...
- dwheeler 6y agoOne cool advantage of having multiple compilers for a language is that you can use one as a check on the other. For example, if you're worried that one of the compilers might be malicious, you can use the other compiler to check on it: https://dwheeler.com/trusting-trust https://dwheeler.com/trusting-trust Even if you're not worried about malicious compilers, you can generate code, compiled it against multiple compilers, and sending inputs and see when they differ in the outputs. This has been used as a fuzzing technique to detect subtle errors in compilers.
- steveklabnik 6y agoYep! This is a very good property, and part of why mrustc is a big deal.
- gbrown_ 6y ago> For example, if you're worried that one of the compilers might be malicious, you can use the other compiler to check on it: https://dwheeler.com/trusting-trust https://dwheeler.com/trusting-trust This still requires the use of a use of trusted compiler though. Comparing two compilers arbitrarily shows if there is consensus, it does not give guarantees about correctness. From the link. In the DDC technique, source code is compiled twice: once with a second (trusted) compiler (using the source code of the compiler’s parent), and then the compiler source code is compiled using the result of the first compilation. If the result is bit-for-bit identical with the untrusted executable, then the source code accurately represents the executable.
- dwheeler 6y agoFirst, I forgot to disclose: I am the author of https://dwheeler.com/trusting-trust https://dwheeler.com/trusting-trust . As discussed in detail in that dissertation, if you are using diverse double compiling to look for malicious compilers, the trusted compiler does not have to be perfect or even non-malicious. The trusted compiler could be malicious itself. The only thing you're trusting is that the trusted compiler does not have the same triggers or payloads as the compiler it is testing. The diverse double compiling check merely determines whether or not the source code matches the executable given certain assumptions. The compiler could still be malicious, but at that point the maliciousness would be revealed in its source code, which makes the revelation of any malicious code much, much easier. You're absolutely right about the general case merely showing consistency, not correctness. I completely agree. But that still is useful. If two compilers agree on something, there is a decent chance that their behavior is correct. If two computers disagree on something, perhaps that is an area where the spec allows disagreement, but if that is not the case then at least one of the compilers is wrong. The check by itself won't tell you whirch one is wrong, but at least it will tell you where to look. In a lot of compiler bugs, having some sample code that causes the problem is the key first step.
- Myrmornis 6y agoThere wouldn't be any surprises, or cognitive dissonance, from using very different paths for debug versus release builds? On a small project, personally I use --release sometimes during development because the compile time doesn't matter that much and the resulting executable is much faster: if I don't use --release I can get a misleading sense of UX during development.
- steveklabnik 6y agoThis already happens a bunch, even with the current setups. It's very natural if you come from a compiled language, and not if you don't. The first step of someone saying "hey why is Rust slow?" is five people replying "did you use --release".
- sfink 6y agoCan confirm. I had a graph traversal program written in Python. I ported it to Rust, and the runtime was identical -- 68.4 seconds, down to the tenth of a second. (Kinda blew my mind -- I had to triple check that I was running and timing what I thought I was!) I had a bit of a crisis of faith. I poked at it a few times over the next week, then finally got on the IRC channel and quickly received the advice mentioned above. Same input, with --release: 6.2 seconds.
- Leherenn 6y agoIt's funny, because I do the exact opposite. As a developer I usually have a pretty powerful machine, and I've found that debug mode is a good way to approximate slow computers, and something that is unbearably slow in debug will bother some users later on.
- runevault 6y agoThis is an interesting idea, but I guess my one question is how much does the slowness of debug relate to HOW it will be slow in release? Since release optimizations can do pretty radical things to the assembly generated it feels like it wouldn't really be apples to apples.
- WalterBright 6y agoThe D programming language has 3 compilers, one with LLVM (LDC) one with GCC (GDC) and one with the Digital Mars back end (DMC). It's great to have all three, as they each have different characteristics in terms of speed, generated code, debug support, platform support, etc. Supporting these three also helps maintain proper semantic separation of code gen from front end.
- stjohnswarts 6y agoHas the D community been growing or shrinking over the past decade or so? Staying relatively the same size?
- stackzero 6y ago+1 enjoyed how accessible this write up was
- brokenbotnet 6y agoReally great.
- tester3 6y agoHow do they ensure that output of both compilers is correct? e.g LLVM output is A, but the new one is B, how do they deal with different results between backends?