10 ms·
The rev.ng decompiler goes open source
- Fnoord 3y agoPrice model: > Very briefly: > The rev.ng framework is fully open source. You can decompile anything you want from the CLI. > The UI will be available in the following forms: > free to use in the cloud for public projects; > available through a subscription in the cloud for private projects; > available at a cost as a fully standalone, fully offline application. In comparison, Hopper costs 100 USD with one year of updates [1]. Ghidra and Radare2 are FOSS and completely free to use, IDA Pro costs a fortune [1] https://www.hopperapp.com/index.html https://www.hopperapp.com/index.html
- eyegor 3y agoBinary ninja is another good option. In my experience it's pretty similar to ida but I find it more user friendly. It just has a lot of well thought out features that make me more productive. I haven't tried hopper but ghidra and radare2 both had a bad dev experience and produced c that didn't "read well". Granted it's been a couple of years since I tried either. Binja is $300 (or $1500 for commercial, both cheaper for students). https://binary.ninja/features https://binary.ninja/features
- aleclm 3y agoStudents shouldn't pay a dime. They are poor. Our view is: the engine is 100% open source. The UI is available for free in the cloud for anyone experimenting, which we define as "I'm OK with leaving the project public". Basically, the decompiler engine is Free Software, extensible and available for automation/scripting, while the UI is available for free for students/researchers and we can make a living out of professionals (i.e., when your company is paying for it).
- halayli 3y agoThey are now offering a free version: https://binary.ninja/2024/02/28/4.0-dorsai.html https://binary.ninja/2024/02/28/4.0-dorsai.html
- Nereuxofficial 3y agoOh that is awesome! I've used the cloud version previously but now that the desktop version is free with some small limitations i think I'll probably use it instead of Ghidra
- felipefar 3y agoI really like licensing models of one-time payments with a pre-defined duration of updates. But I wonder how they enforce it while not making internet access a requirement for the app.
- 8organicbits 3y agoI've been planning to use a non-enforcement model for a future project. Some users will always pay, because of corporate policy or ethics. Some will never pay and will reverse engineer out any software license checks. Asking the user if they have a license keeps the honest ones honest and permits ad-hoc free trials, emergency use, and other reasonable "unlicensed use".
- userbinator 3y agoSome will never pay and will reverse engineer out any software license checks. For a long time (and might still be; not paying much attention anymore), it was a "rite of passage" in the scene to crack IDA... using itself.
- mrexodia 3y agoIt never was a “rite of passage”, because removing IDA’s license checks has always been trivial…
- deleted 3y ago[deleted]
- dvzk 3y agoDecompilation is often the least important (and least reliable) part of IDA/Ghidra, so comparing the two is unfair. That said, the scene is perpetually starved for good C decompilers, so more attempts are always exciting.
- saagarjha 3y agoI hear this a lot and in my experience people who Ghidra or IDA and don’t use the decompiler are exceptionally rare. Why would you suffer that when you can use something else for what you actually want?
- dvzk 3y agoI didn't say I never use it, just that it's not always the core feature. This will depend heavily on your field, but in my past work, the features that were way more essential are: scripting (+ IR lifting), xrefs, CFGs, labels/notes (in a persistent DB). In my experience decompilers will totally ignore or fail on certain types of malicious code, so they mainly exist to assist disassembly analysis. And for that purpose, they save us an incredible amount of human hours.
- aleclm 3y agoFor scripting, our approach is to give you access to the project file (just a YAML file), and you can make changes from any scripting language you want. Everything the user can customize is in there, all the rest is deterministically produced from that file. I really disliked the fact that you usually need to buy into the version of Python that $TOOL requires you to use, or the fact itself that you need to use a specific language. Can parse YAML? You're mostly done. The "project file" is what we call the model: https://docs.rev.ng/user-manual/model-tutorial/ https://docs.rev.ng/user-manual/model-tutorial/ For xrefs, CFG and the rest: we have all of that in the UI, but we also produce them in a rich way. For instance, when we emit disassembly and decompiled code, we actually emit plain text + HTML-like markup to provide metainformation for navigation (basically, xrefs) and highlighting. So you can use all that from any language that can parse HTML/XML. It's called PTML: https://docs.rev.ng/references/ptml/ https://docs.rev.ng/references/ptml/ For lifting: we use LLVM IR as our internal representation. This means that: 1) you don't have to learn an IR that no one else uses, 2) you can use off the shelf tools (e.g., KLEE for symbolic execution) but you can also use all the standard LLVM optimizations and analyses and 3) you can recompile it, but we're not into the binary translation business anymore.
- nextos 3y agoA cool company fueled by one of the best PLT books out there: https://link.springer.com/book/10.1007/978-3-662-03811-6 https://link.springer.com/book/10.1007/978-3-662-03811-6 "He also met a partner in crime, Pietro. Romantically enough, he met him thanks to a book which will turn out to be foundational for company." https://rev.ng/about https://rev.ng/about Congrats on the launch.
- aleclm 3y agoAbout the book, here's the full story: I was getting into compilers, but I was really struggling with the theory, the most famous books weren't doing it for me, and I felt really down. Then I find this book, which seems very dense, but clear. So I ask my advisor if I could buy it and goes like "well, first check out the university library". I check it out and there's a copy, but... it's taken. Working in the only group that was doing research on compilers I'm like "who dares do compilers stuff out of our group!?". I go to the library: Me: who has the book? Library guy: can't tell you, privacy reasons. Me: what's the third letter of its surname? Library guy: Z Me: what's the second letter of its name? Library: I Me: thanks. I go here: https://www.deib.polimi.it/ita/personale-lista-alfabetica https://www.deib.polimi.it/ita/personale-lista-alfabetica I found him. Fast forward, we become friends and we start the company together. > Congrats on the launch. Thanks! It was a lot of work.
- albertzeyer 3y agoChecking the team about: https://rev.ng/about https://rev.ng/about And looking at the code contributions: https://github.com/revng/revng/graphs/contributors https://github.com/revng/revng/graphs/contributors Isn't it a bit weird that the CEO (aleclearmind) has most commits, even much more than the CTO (pfez)? I often hear the complaints from other CEOs that they don't really find any time anymore to code... Even the CTO usually is more on the managing side and less active in actual coding. Anyway, if this works, then I guess it's a lot of fun for them. Edit Ah right, I didn't check the timeline.
- zote 3y agoThe CTO has more recent commits, aleclearmind's commits drop to 0 after 2020 so maybe they also have a hard time getting to code.
- aleclm 3y agoThe CTO mostly works on the backend of the decompiler, revng-c, which we just released: https://github.com/revng/revng-c/commits/develop/ https://github.com/revng/revng-c/commits/develop/ Eventually we'll merge the two repos. Also, I develop stuff every day. For some reason GitHub is not picking up my user correctly. > Anyway, if this works, then I guess it's a lot of fun for them. It is!
- albertzeyer 3y agoI wonder a bit about the downvotes. I didn't mean this as a criticism or so in any way. In fact, I like this very much. I just found this interesting and unlike what I saw elsewhere. So the downvotes are because this is not interesting or not unusual?
- halayli 3y agoyour observation was spot on and your question was answered by the ceo. People on hn can be oversensitive.
- londons_explore 3y agoIdea: automatically name variables and members of structs based on how code interacts with them. Eg. The next pointer in a linked list should be easy to identify as 'next'. That would be done by downloading all of GitHub, then seeing what variables in GitHub code have the most similar layouts and interactions, and then if the confidence is high enough, using those names.
- qweqwe14 3y agoSort of like GitHub Copilot but for reversing?
- aleclm 3y agoIn the past we were thinking to do something like this by hand. For instance, we detect induction variables, we could rename them into `i`. However, nowadays, it seems pretty obvious that the right way to do this things is using LLMs. This said, at this stage, we see ourselves as people building robust infrastructure. Once the infrastructure is there, using some off the shelf model to rename things or add comments is relatively easy. Basically: we do the hard decompilation work that needs 100% accuracy, and then we can adopt LLMs for things that are OK to be approximate such as names, comments and the like. Anyway, writing a script that renames stuff is pretty easy. Check out the docs: https://docs.rev.ng/user-manual/model-tutorial/ https://docs.rev.ng/user-manual/model-tutorial/
- londons_explore 3y agoIf an LLM is used, it's unclear how to best do it. One could try to train ones own LLM from scratch, using an encoder-decoder (translation - aka seq2seq) architecture trying to predict the correct variable name given the decompiled output. One could try to use something like GPT-4 with a carefully designed prompt "Given this datastructure, what might be the name for this field?" One could try to use something pretrained like llama, but then finetune it based on hundreds of thousands of compiled and decompiled programs.
- Eisenstein 3y ago
- yakkityyak 3y agoI hope collaborative workflows get a lot of attention. I haven't used IDA teams or anything, but a reverse engineering experience that felt as frictionless as Google Docs would be amazing.
- aleclm 3y agoThat's our goal. We used to use QtCreator as a basis for the UI, terrible idea. Then we switched to VSCode, which happens to be able to run in the browser. So we added some magic kubernetes sauce and voilà, you got the cloud decompiler with exactly the same user experience as the fully standalone one. We still need to perform some QA on collaboration, but basically works. One daemon, many clients. Very simple architecture. I think we got inspiration to do this from a CTF where we were doing "collaboration" using IDA with multiple windows on a X session on a server with multiple cursors. Very cursed, but effective.
- fwr00t 3y agoSeems exciting. I'm keen to try the fully standalone version. Is there any news about tentative pricing? Hopefully its affordable enough for hobbyist as well.
- JonChesterfield 3y agoAlways pleased to see more binary hacking tools. A load of overly-precise suggestions on the chosen packaging format follows because I might want to use this tool myself :) > `source ./environment` That's a bad omen. I downloaded the tar to find it does indeed set a bunch of environment variables including PATH, though thankfully not LD_LIBRARY_PATH. Mostly prefixed "HARD_" which is maybe unique (REVNG would be a more obvious choice, colliding with existing environment variables is a bad thing). It sets `AWS_EC2_METADATA_DISABLED="true"` which won't break me (I don't use AWS) but in general seems dubious. export RPATH_PLACEHOLDER="////////////////////////////////////////////////$ORCHESTRA_ROOT" export HARD_FLAGS_CXX_CLANG="-stdlib=libc++" ... "-Wl,-rpath,$RPATH_PLACEHOLDER/lib ... This is suboptimal. The very long PATH setting with mingw32 and gentoo and mips strings in it also looks very fragile. I usually bail when the running instructions include "now mangle your environment variables" because that step is really strongly correlated with programs that don't work properly on my non-ubuntu system. Wiring your application control flow through the launching environment introduces a lot of failure modes - it's not as convenient as it first appears. Very like global variables. Clang will burn a lot of this stuff in as defaults when you build it if you ask, e.g. `-DCLANG_DEFAULT_CXX_STDLIB=libc++` would remove the stdlib setting environment variable. DEFAULT_SYSROOT is useful too. Using rpath means you're vulnerable to someone running this script with LD_LIBRARY_PATH set as the environment variable will override your DT_RUNPATH setting in the binaries. The background on this is aggravating. Abbreviating here, '-Wl,rpath' no longer means rpath, it means 'runpath' which is a similar but much less useful construct. The badly documented invocation you probably want is `-Wl,rpath -Wl,--disable-new-dtags` to set rpath instead of set runpath, at which point the loader will ignore LD_LIBRARY_PATH when looking for libraries. There's a good chance you can completely remove the environment mangling through a combination of setting different flags when building clang, static linking and embedding binaries in other binaries. Related, your clang-16 binary is dynamically linked. As in it goes looking for things like libLLVMAArch64CodeGen.so.16 at runtime. A lot of failure modes can be removed by LLVM_BUILD_STATIC=ON. E.g. if I run your dynamically linked clang with a module based HPC toolchain active, your compiler will pick up the libraries from the HPC toolchain and it'll have a bad time. The tools are all linked against glibc as well, pros and cons to that. Tools are also linked against libc++.so, which is linked against libc++abi.so and so forth. Worth considering static libc++, but even if you decline that, libc++abi and libunwind can and probably should be statically linked into the libc++. The above rpath rant? Runpath isn't transitive so dynamic libaries finding other dynamic libraries using runpath (the one you get when you ask for rpath) works really poorly. Context for there being so many suggestions above - I am completely out of patience with distributing dynamically linked programs on Linux. I don't want a stray environment variable from some program that had `source ourhack` in the readme or a "module system" to reach into my application and rewire what libraries it calls at runtime as the user experience and subsequent bug report overhead is terrible. Static linking is really good in comparison. Thanks again for shipping, and I hope some of the above feedback is helpful!
- dark-star 3y agoIt doesn't work with my ELF file: [orchestra] [darkstar@shiina revng]$ ./revng artifact --analyze --progress decompile-to-single-file ../maytag.ko [=======================================] 100% 0.57s Analysis list revng-initial-auto-analysis (5): import-binary [===================> ] 50% 0.57s Run analyses lists (2): revng-initial-auto-analysis [=========> ] 25% 0.57s revng-artifact (2): Run analyses Only ELF executables and ELF dynamic libraries are supported [orchestra] [darkstar@shiina revng]$ file ../maytag.ko ../maytag.ko: ELF 64-bit LSB relocatable, x86-64, version 1 (FreeBSD), not stripped Does it not support FreeBSD binaries? Edit: Ah I missed that it doesn't support kernel modules, probably has nothing to do with FreeBSD but the fact that this is not a simple executable
- aleclm 3y agoCan you open an issue on GitHub and attach the binary? I don't think it should be too hard to load that.
- dark-star 3y agocan I somehow share the binary privately? It's a proprietary module that I probably shouldn't share publically (also it's ... rather large) I opened issue #366 for it already
- costco 3y agoCongrats. Do you have any regrets about outsourcing lifting to the QEMU TCG or has it worked well?
- aleclm 3y agoThanks! It has been working very well. Two regrets: 1. Not rebasing our fork of QEMU for years has put us in a bad spot. But just today a member of our team managed to lift stuff with the latest QEMU. And he has also been able to lift Qualcomm Hexagon code, for which we helped to add support in QEMU. Eventually we'll be the first proper Hexagon decompiler :) 2. Focusing too much on QEMU led our frontend to be tightly coupled with QEMU. It will now take some effort to enable support for additional frontends, non-QEMU based. But not impossible: our idea is to let user add support for a new architecture by defining, in C, a struct for the CPU state and a bunch of functions acting on it. That's it. No need to learn any internal representation. tl;dr QEMU was a great choice, it worked so well that we didn't work on that part of the codebase for too much time and now there's some technical debt there. But we're addressing it.
- flexagoon 3y agoAre there any plans to support type inference? It seems like it currently shows all variables as generic64_t. Would be nice to automatically detect their types like Ghidra does (albeit sometimes incorrectly)
- aleclm 3y agoYes. Roadmap item: https://rev.ng/roadmap#feature-798 https://rev.ng/roadmap#feature-798 Design pad: https://pad.rev.ng/s/eDHi2PUoP# https://pad.rev.ng/s/eDHi2PUoP#