5 ms·
As usual, Twitter is impressed by this, but I'm very skeptical, the chance of it breaking your program is pretty high. The thing that makes optimizations so har
by chad1n 2y ago
As usual, Twitter is impressed by this, but I'm very skeptical, the chance of it breaking your program is pretty high. The thing that makes optimizations so hard to make is that they have to match the behavior without optimizations (unless you have UBs), which is something that LLMs probably will struggle with since they can't exactly understand the code and execution tree.
- ramon156 2y agoLet's be real, at least 40% of those comments are bots
- extheat 2y agoPeople simply have no idea what they're talking about. It's just jumping on to the latest hype train. My first impression here was per the name that it was actually some sort of compiler in it of itself--ie programming language in and pure machine code or some other IR out. It's got bits and pieces of that here and there but that's not what it really is at all. It's more of a predictive engine for an optimizer and not a very generalized one for that. What would be more interesting is training a large model on pure (code, assembly) pairs like a normal translation task. Presumably a very generalized model would be good at even doing the inverse: given some assembly, write code that will produce the given assembly. Unlike human language there is a finite set of possible correct answers here and you have the convenience of being able to generate synthetic data for cheap. I think optimizations would arise as a natural side effect this way: if there's multiple trees of possible generations (like choosing between logits in an LLM) you could try different branches to see what's smaller in terms of byte code or faster in terms of execution.
- quonn 2y ago> Presumably a very generalized model would be good at even doing the inverse: given some assembly, write code that will produce the given assembly. ChatGPT does this, unreliably.
- hughleat 2y agoIt can emulate the compiler (IR + passes -> IR or ASM). > What would be more interesting is training a large model on pure (code, assembly) pairs like a normal translation task. It is that. > Presumably a very generalized model would be good at even doing the inverse: given some assembly, write code that will produce the given assembly. Is has been trained to disassemble. It is much, much better than other models at that.
- solarexplorer 2y agoIf I understand correctly, the AI is only choosing the optimization passes and their relative order. Each individual optimization step would still be designed and verified manually, and maybe even proven to be correct mathematically.
- boomanaiden154 2y agoRight, it's only solving phase ordering. In practice though, correctness even over ordering of hand-written passes is difficult. Within the paper they describe a methodology to evaluate phase orderings against a small test set as a smoke test for correctness (PassListEval) and observe that ~10% of the phase orderings result in assertion failures/compiler crashes/correctness issues. You will end up with a lot more correctness issues adjusting phase orderings like this than you would using one of the more battle-tested default optimization pipelines. Correctness in a production compiler is a pretty hard problem.
- hughleat 2y agoThere are two models. - foundation model is pretrained on asm and ir. Then it is trained to emulate the compiler (ir + passes -> ir or asm) - ftd model is fine tuned for solving phase ordering and disassembling FTD is there to demo capabilities. We hope people will fine tune for other optimisations. It will be much, much cheaper than starting from scratch. Yep, correctness in compilers is a pain. Auto-tuning is a very easy way to break a compiler.
- Lockal 2y agoAs this LLM operates on LLVM intermediate representation language, the result can be fed into https://alive2.llvm.org/ce/ https://alive2.llvm.org/ce/ and formally verified. For those who don't know what to print there: here is an example of C++ spaceship operator: https://alive2.llvm.org/ce/z/YJPr84 https://alive2.llvm.org/ce/z/YJPr84 (try to replace -1 with -2 there to break). This is kind of a Swiss knife for LLVM developers, they often start optimizations with this tool. What they missed is to mention verification (they probably don't know about alive2) and comparison with other compilers. It is very likely that LLM Compiler "learned" from GCC and with huge computational effort simply generates what GCC can do out of the box.
- boomanaiden154 2y agoI'm reasonably certain the authors are aware of alive2. The problem with using alive2 to verify LLM based compilation is that alive2 isn't really designed for that. It's an amazing tool for catching correctness issues in LLVM, but it's expensive to run and will time out reasonably often, especially on cases involving floating point. It's explicitly designed to minimize the rate of false-positive correctness issues to serve the primary purpose of alerting compiler developers to correctness issues that need to be fixed.
- hughleat 2y agoYep, we tried it :-) These were exactly the problems we had with it.
- boomanaiden154 2y agoI'm not sure it's likely that the LLM here learned from gcc. The size optimization work here is focused on learning phase orderings for LLVM passes/the LLVM pipeline, which wouldn't be at all applicable to gcc. Additionally, they train approximately half on assembly and half on LLVM-IR. They don't talk much about how they generate the dataset other than that they generated it from the CodeLlama dataset, but I would guess they compile as much code as they can into LLVM-IR and then just lower that into assembly, leaving gcc out of the loop completely for the vast majority of the compiler specific training.
- bbor 2y agoAFAIK this is a heuristic, not a category. The underlying grammar would be preserved. Personally I thought we were way too close to perfect to make meaningful progress on compilation, but that’s probably just naïveté
- boomanaiden154 2y agoI would not say we are anywhere close to perfect in compilation. Even just looking at inlining for size, there are multiple recent studies showing ~10+% improvement (https://dl.acm.org/doi/abs/10.1145/3503222.3507744 https://dl.acm.org/doi/abs/10.1145/3503222.3507744, https://arxiv.org/abs/2101.04808 https://arxiv.org/abs/2101.04808). There is a massive amount of headroom, and even tiny bits still matter as ~0.5% gains on code size, or especially performance, can be huge.
- cec 2y agoHey! The idea isn't to replace the compiler with an LLM, the tech is not there yet. Where we see value is in using these models to guide an existing compiler. E.g. orchestrating optimization passes. That way the LLM won't break your code, nor will the compiler (to the extent that your compiler is free from bugs, which can tricky to detect - cf Sec 3.1 of our paper).
- verditelabs 2y agoI've done some similar LLM compiler work, obviously not on Meta's scale, teaching an LLM to do optimization by feeding an encoder/decoder pairs of -O0 and -O3 code and even on my small scale I managed to get the LLM to spit out the correct optimization every once and a while. I think there's a lot of value in LLM compilers to specifically be used for superoptimization where you can generate many possible optimizations, verify the correctness, and pick the most optimal one. I'm excited to see where y'all go with this.
- viraptor 2y agoThank you for freeing me from one of my to-do projects. I wanted to do a similar autoencoder with optimisations. Did you write about it anywhere? I'd love to read the details.
- verditelabs 2y agoNo writeup, but the code is here: https://github.com/SuperOptimizer/supercompiler https://github.com/SuperOptimizer/supercompiler There's code there to generate unoptimized / optimized pairs via C generators like yarpgen and csmith, then compile, train, inference, and disassemble the results
- hughleat 2y agoYes! An AI building a compiler by learning from a super-optimiser is something I have wanted to do for a while now :-)
- moffkalast 2y ago
- namaria 2y agoThis feels like going insane honestly. It's like reading that people are super excited about using bouncing castles to mix concrete.