7 ms·
A Complete Guide to LLVM for Programming Language Creators
- TazeTSchnitzel 6y ago> LLVM IR looks like a more readable form of assembly I think assembly is more readable than the dumped LLVM IR output of a real compiler… Anyway, that was a good overview of how to build simple LLVM IR. :)
- mhh__ 6y agoThe IR is pretty messy but if you can see through the clutter the IR is much easier to read the flow of. There's a reason why we use SSA (although prior to mem2reg running you don't quite have it, but still)
- mrathi12 6y agoYes, exactly. Clang emitted IR especially has a lot of (C/C++)-specific junk. If you look past that clutter, it's not too bad. I think the best way to learn to read IR is to look at super-minimal examples, and then you'll be able to tell which parts of larger IR files are relevant.
- RicardoLuis0 6y agoI'd say readability of dumped IR would be on par with dumped assembly from a compiler for readability, but hand-written LLVM IR has the potential to be quite a lot more readable than hand-written assembly
- mraza007 6y agoReally loved reading it. It made it so much easier for me understand LLVM. Just wanted to appreciate you for writing this awesome guide
- mrathi12 6y agoHi, author of the post here. Thank you very much for your kind words!
- ufo 6y agoA question for the LLVM experts here: In the past when I've looked at LLVM from a distance, the biggest stumbling block I found were that it's written in C++ , which isn't the language I'm using for my frontend. How important is the C++ API in practice? Are the bindings for other languages usable? Is it possible to have my frontend emit LLVM IR in a text format, similarly to how you can feed assembly language source code to an asssembler? Or should one really just bite the bullet and use C++ to generate the IR? I noticed that the compiler in this tutorial has a frontend in Ocaml and a backend in C++, with the communication between them being done via protobufs.
- mhh__ 6y agoI would avoid any text, but LLVM has mature bindings in ocaml and Haskell, for example. The textual representation isn't stable IIRC, and it adds a step in between you and your already lumbering backend. Ultimately the C++ API isn't too difficult to use but LLVM mandates a fairly hardcore level of C++ knowledge to play with it's internals. Quick Tip: if you're thinking "Holy shit how do I get from [complicated] all the way down to IR instructions" lower the big thing to a something simpler so you can reuse the code to generate IR from that - for example, a foreach loop is expressible as a for loop under the hood, now you only have to be compile for loops. This would usually be done in the AST itself.
- ritter2a 6y agoRegarding interface stability: Indeed, the textual representation is not stable, things like added types in the representation of some instructions can happen when upgrading to a new version. However, to be entirely honest, in the last few years of updating LLVM-based research tools to newer LLVM versions, changes in the C++ API that required me to (sometimes just slightly) change my code happened a lot more often than changes in the textual representation...
- wallnuss 6y agoJulia has an interesting split here, it does the lowering into SSA from in pure Julia and then has a codegen steps that translates the SSA from into LLVM IR, but for that second step we do use the C++ API. We have very robust bindings to the C-API, but it forever feels just a bit incomplete and less cared for. The C-API is very stable, whereas the C++ API does change quite a bit.
- deleted 6y ago[deleted]
- finiteloop 6y agoAs someone who writes a lot of toy languages, I made this scaffolding for a LLVM-based compiler: https://github.com/finiteloop/compiler https://github.com/finiteloop/compiler It uses Bison and Flex for parsing and lexing unlike this post, but may be a useful starting point for those building their own toy languages.
- dbcurtis 6y agoDoesn’t the LLVM API go through breaking changes fairly often? How do you track all that and keep your sanity?
- finiteloop 6y agoEveryone says that but I have not experienced it. In practice, I think this impacts backend extension developers more than people targeting LLVM IR. My experience covers version 7-11, but perhaps it used to be worse?
- vnorilo 6y agoI've been updating a small codegen since the 3.x days. There have been many API changes with major versions. That said, I always found it pretty easy to implement the changes (something like a workday for my ~10k LLVM-interfacing LoC), and the changes tend to be such that once you get it to compile again, it just works as before. I think LLVM is an excellent demonstration of how to design and implement big systems well in C++, and as a part of that, they also do breaking changes pretty well.
- exDM69 6y ago> Doesn’t the LLVM API go through breaking changes fairly often? How do you track all that and keep your sanity? No, not very big ones. But they give no stability promises. If you're using a language other than C++ via API bindings, you will experience some churn going from version to version as the bindings need to be updated too.
- chrisaycock 6y agoReusable compiler components are really helpful. I made this example for parsing with ANTLR: https://github.com/empirical-soft/calculANTLR https://github.com/empirical-soft/calculANTLR It uses a C++ port of CPython's ASDL to define the AST.
- HexDecOctBin 6y agoIs there a similar resource for liblldb? The only documentation I could find was the Doxygen-based one, and it doesn't provide enough hints to a newcomer to know where to get started if I want to write a debugger frontend.
- CGamesPlay 6y agoIt’s possible I missed this from the article, since I’m not actually following along, but why do I need to create the custom C++ adapter? The article mentions both ll bc formats, but I don’t see how they are used? If I wanted to target LLVM, why wouldn’t I emit ll or bc files?
- tylersmith 6y agoWhat do you mean by C++ adapter? All the c++ in this post is building the LL files in memory, which can them be persisted to disk or wherever.
- CGamesPlay 6y agoAh I didn’t understand that aspect. So we write the parser in ocaml, then serialize it to protobuf, then consume that from C++ so we can call LLVM’s API to write the expected file format? That still seems needlessly complicated. Is the file format simply too arcane?especially given [0]. [0] https://news.ycombinator.com/item?id=25540637 https://news.ycombinator.com/item?id=25540637
- tylersmith 6y agoI see, in this case the author's compiler was written in c++, but you can use ocaml instead to build the IR. There is even a version of their great tutorial series using ocaml you can check out. All the concepts he outlines are the same, just convert the c++ bits to the corresponding ocaml version of the IRBuilder. Then you can go from your language to llvm IR all in ocaml. https://llvm.org/docs/tutorial/OCamlLangImpl1.html https://llvm.org/docs/tutorial/OCamlLangImpl1.html
- CGamesPlay 6y agoAh, I understand now. Thanks for helping!
- kevindeasis 6y agoThanks for this, I used to have other resources too about 3 years ago when I was making my own toy blockchain project that takes in a programming language to create smart contracts let me see if i can still find those
- ithkuil 6y agoDoes anybody know a good resource like this but for writing backends for a new cpu ISA?
- RealityVoid 6y agoI would like to know this as well. I mostly work with Renesas RH850 cores and was interested to be able to compile rust for them, but there is no v850 backend, sadly. Was considering fiddling with it, but couldn't figure out how. I also hate the proprietary compilers for this arch, since they are sooo slooow (maybe for good reason, but I have no comparison point) and their licensing method just annoys me.
- Rochus 6y agohttps://llvm.org/docs/WritingAnLLVMBackend.html https://llvm.org/docs/WritingAnLLVMBackend.html https://jonathan2251.github.io/lbd/ https://jonathan2251.github.io/lbd/
- nevi-me 6y agoI'm playing around trying to create a JIT execution engine for Rust code. I was struggling a bit with the codegen part, so the description of the function structure helped make more sense of what I'm trying to do. Thanks
- mrathi12 6y agoWhile we're here, let's not forget about the incredible Kaleidoscope tutorial. They really helped me get a grip with LLVM. https://llvm.org/docs/tutorial/index.html https://llvm.org/docs/tutorial/index.html
- mannykannot 6y agoThis is most welcome, as I recently started looking at LLVM, and I was not finding it easy to get a clear picture of the project's organization. I see that your username resembles that of the entity hosting this guide - I'm guessing that's not a coincidence!
- jooz 6y agoWhen I studied compilers back in the university, the subject consist in reading understanding and putting in practice the 'dragon book' (not the full book but a big part of it). We all used flex+yacc and a simplified ASM that was interpreted by some educational software which name I cant remember. The project for the year was to implement a basic compiler language: included if/for loops, function calls, recursivity ... basic stuff but enough as starting point to build something. Trying to get into LLVM, Id love to find an example done in flex+yacc vs the same done in LLVM.
- ufo 6y agoFlex+yacc and llvm are for separate steps in the process. The former are for parsing the source code into an AST and the latter for generating executable code from the AST. That said, somewhere else on this comment thread user "finiteloop" posted the scaffolding for a toy compiler using yacc+llvm. You could check that out.
- jtsiskin 6y agoLLVM: used to be “Low Level Virtual Machine”, but the acronym no longer applies IR: intermediate representation