3 ms·
> Running O2 optimizations sounds great in theory, but is often not too helpful in practice I cannot stress how important it is to have at your disposal an ali
by aleclm 3y ago
> Running O2 optimizations sounds great in theory, but is often not too helpful in practice
I cannot stress how important it is to have at your disposal an alias analysis framework and analyses such as LazyValueInfo and ScalarEvolution. In fact, some whole key features of rev.ng are possible thanks to them.
Either one (thinks) they're not need them or you have to reinvent the wheel. People think binary analysis is something special, it's not, just compilers. And there's really really a lot to learn from the design of LLVM.
Add with carry doesn't sound like a gigantic problem to me, given what you get in return. About vector instructions, a lot of autovectorization takes place in the mid-end and LLVM supports vector types, but I'm not sure if you're talking about them in input or in output.
If you really want to roll your own thing, one could at least use MLIR, where you can define your own operations but you reuse MLIR infrastructure, pass manager, DCE and much more. However, you'd lose the above mentioned passes in the early stages of the pipeline is too much of a loss.
As mentioned, we started facing the limitations of LLVM when you want sophisticated types (e.g., unions) and more and more high level concepts while you get close to C. That's why we're rolling our MLIR dialect, clift.
> LLVM-IR is fairly heavy-weight and not particularly efficient
I guess it depends what's the comparison, but in our experience, it scales pretty pretty well.
But overall... no one wants to learn your custom IR, there's much more incentive in learning an (the most?) established IR, in particular for a lifter.
If you use LLVM IR you get AddressSanitizer, CoverageSanitizer, KLEE, polly, libFuzzer, bindings, documentation and so much more.
- aengelke 3y ago> People think binary analysis is something special, it's not, just compilers. I (conceptually) agree! But there's different sorts of compilers, and they have different requirements, which (should) materialize in the code representation. An offline C/C++ compiler has different requirements than a JIT JavaScript compiler, and so their architecture is very different. The IR should be express the relevant operations/data structures well and should easily support the analyses/transformations. If LLVM is a good fit for you, that's great. But in many cases, it's not. I've been bitten from being warned about not writing my own IR (effort concerns), and in hindsight, it would've been the much easier way, both for low-level and high-level transformations. IMHO, people should not be afraid to roll their own IR and make it a good fit for their use case. > one could at least use MLIR MLIR is a nice idea (and great for selling/publishing). I find the implementation to be somewhat lacking and constraining, difficult to work with, and not particularly efficient. > no one wants to learn your custom IR That's a problem, but there's no silver bullet.
- aleclm 3y agoI see your position, but I've to say that in 8 years of rev.ng we revised many many design choices we initially made, but not using LLVM was not one of them, so I can't agree. Rolling your own IR is not hard, in fact, it's very easy (and sometimes we roll temporary IRs for specific purposes), but it's also true that most IRs look more or less the same. QEMU's tiny code and LLVM IR, for instance, are quite similar. What matters is the ecosystem around them, and the ecosystem around LLVM is really great. Anyway, every Friday morning 11:00 CEST we have the rev.ng hour, our weekly internal technical meeting. It would be interesting to discuss what limitations you hit with LLVM that we did not hit and your concerns about efficiency of LLVM and MLIR. We used to stream the rev.ng hour and we're in the process of going back to publish some talks every now and then. If you're interested, drop me an e-mail! :)