3 ms·
Cool projects, but let me make two observations: 1. Writing a binary analysis tool is an enormous task. Before getting to the useful and new part, you need a l
by aleclm 3y ago
Cool projects, but let me make two observations:
1. Writing a binary analysis tool is an enormous task. Before getting to the useful and new part, you need a lot of infrastructure (e.g., a scalable intermediate representation). Rolling your own infrastructure is not a great idea IMO, there are great compiler frameworks (LLVM, but GCC would work too) that provide you with a lot of tools to start with.
We can literally run the LLVM -O2 optimization pipeline on our IR.
2. In binary analysis there's a trend to use SMT solvers. We absolutely do not, by design. They simply do not scale.
Binary analysis theory is really just compiler theory. Have you ever seen a mainstream compiler use SMT solvers for the standard compilation pipeline? No, it does not scale.
Depending on what you need to do, you go the old fashioned way and write a data-flow analysis that does what you need, instead of using an SMT solver that can solve much more complex problems but does not scale.
At rev.ng, we're in love with Monotone Frameworks. If your analysis fits in the framework, it will run in linear time. It's more difficult than going "Z3 solve it for me", but it scales great.
Suggested reading: https://link.springer.com/book/10.1007/978-3-662-03811-6 https://link.springer.com/book/10.1007/978-3-662-03811-6
Don't get me wrong, there are use cases where using an SMT solver makes sense, for instance looking for bugs, but SMT solvers should be a last resource measure for tasks where you're OK with saying "OK, I didn't find anything in 10 minutes, let's bail out". But that's not the case for simply lifting and even decompiling a binary.
This said, I haven't been following closely the developments of rizin, but my comments are more on the general trend in binary analysis that I see.
- xvilka 3y ago> compiler frameworks (LLVM, but GCC would work too) that provide you with a lot of tools to start with. While I agree with this, the binary code to LLVM IR uplifting loses a lot of context and semantics information because LLVM IR was designed to do precisely the opposite. Moreover, the increasing popularity of using "middle level" IRs in LLVM-based compilers, like MIR and Swift IR, makes the "distance" between code and the IR representation even bigger. This is why IRs specifically designed to tackle RE tasks would perform better (RzIL and any relatively modern intermediate representation). > They simply do not scale. I agree with this notion, but you don't need to "solve" a whole program; it's impossible for any real-world software; you can selectively pick places to do so. Currently, existing binary analysis software already does that, but usually in a more "manual" way, e.g., by emulating pieces of the code sparsely to figure out indirect jumps and so on.
- aleclm 3y ago> the binary code to LLVM IR uplifting loses a lot of context Losing context is good in order to ensure you properly decoupled the frontend from the rest of the pipeline. We don't even keep track of what a "call" instruction is, we re-detect it on the LLVM IR. One reason you may want to preserve context is to let the user know where a specific piece of lifted code originated from. In order to preserve this information, we exploit LLVM's debugging metadata and it works pretty well. There's some loss there, but LLVM transformations strive to preserve it. After all, imagine you have `add rax, 4; add rax, 4`, you'll want to optimize it to a +8 and you'll either have to decide if you want to associate your +8 operation with the first or the second instruction. > the binary code to LLVM IR uplifting loses a lot of [...] semantics information Not sure what you mean here, we use QEMU as a lifter and that's very accurate in terms of semantics. I'm not sure what MIR and Swift IR have to do with the discussion, those are higher level IRs for specific languages. LLVM is rather low level and it's language agnostic. However, for going beyond lifting, i.e., decompilation, it's true that LLVM shows some significant limitations. That's why we're rolling our own MLIR dialect, but we can still benefit of all the MLIR/LLVM infrastructure, optimizations and analyses. We're not starting from scratch. > emulating pieces of the code sparsely to figure out indirect jumps and so on It's hard to emulate without starting from the beginning. Maybe you're thinking about symbolic execution? In any case, rev.ng does not emulate and does not do any symbolic execution: we have a data-flow analysis that detects destinations of indirect jumps and it's pretty scalable and effective. Example of things we handle: https://github.com/revng/revng-qa/blob/master/share/revng/test/tests/analysis/CollectCFG/arm/switch-disjoint-ranges.S https://github.com/revng/revng-qa/blob/master/share/revng/te...
- aengelke 3y ago> Rolling your own infrastructure is not a great idea IMO For the start, I'd agree, but as soon as you want to do non-trivial things, I don't think that standard compiler IRs (e.g., LLVM) are a good way for binary analysis. (POV: been there, done that.) The time investment in fighting the existing IR quickly becomes larger than just writing a custom IR, which is not _that_ much effort. Running O2 optimizations sounds great in theory, but is often not too helpful in practice. LLVM is clearly intended to provide a (not too large) set of semantics that can be mapped easily to different architectures. This is great for compilation, but not for the other direction, as even simple things like add-with-carry end up being multiple non-trivial instructions which aren't handled by existing optimization patterns. The same goes for vector instructions, where many patterns are encoded in the back-ends and not in the IR passes. So you end up writing custom passes fairly quickly, but still fight against an IR that doesn't allow you to express the things you'd like to have. In the long run, a custom IR allows for a much more idiomatic code representation and more flexible analyses; not to mention that LLVM-IR is fairly heavy-weight and not particularly efficient. If I were to write another binary translator (I wrote Instrew/Rellume based on LLVM), I would definitely not choose LLVM. I completely agree on the SMT part. I experimented with SMT, but standard data flow analysis suffice 90% of the time and are much more efficient and predictable.
- aleclm 3y ago> Running O2 optimizations sounds great in theory, but is often not too helpful in practice I cannot stress how important it is to have at your disposal an alias analysis framework and analyses such as LazyValueInfo and ScalarEvolution. In fact, some whole key features of rev.ng are possible thanks to them. Either one (thinks) they're not need them or you have to reinvent the wheel. People think binary analysis is something special, it's not, just compilers. And there's really really a lot to learn from the design of LLVM. Add with carry doesn't sound like a gigantic problem to me, given what you get in return. About vector instructions, a lot of autovectorization takes place in the mid-end and LLVM supports vector types, but I'm not sure if you're talking about them in input or in output. If you really want to roll your own thing, one could at least use MLIR, where you can define your own operations but you reuse MLIR infrastructure, pass manager, DCE and much more. However, you'd lose the above mentioned passes in the early stages of the pipeline is too much of a loss. As mentioned, we started facing the limitations of LLVM when you want sophisticated types (e.g., unions) and more and more high level concepts while you get close to C. That's why we're rolling our MLIR dialect, clift. > LLVM-IR is fairly heavy-weight and not particularly efficient I guess it depends what's the comparison, but in our experience, it scales pretty pretty well. But overall... no one wants to learn your custom IR, there's much more incentive in learning an (the most?) established IR, in particular for a lifter. If you use LLVM IR you get AddressSanitizer, CoverageSanitizer, KLEE, polly, libFuzzer, bindings, documentation and so much more.