5 ms·
I'll bite! What wheel is this?
by nssnsjsjsjs 1y ago
I'll bite! What wheel is this?
- Archelaos 1y agoProbably what he links to on his profile page.
- weaksauce 1y agogot to imagine it is this https://github.com/boricj/ghidra-delinker-extension https://github.com/boricj/ghidra-delinker-extension
- boricj 1y agoThis wheel: https://github.com/boricj/ghidra-delinker-extension https://github.com/boricj/ghidra-delinker-extension It's a Ghidra extension that can export relocatable object files from any program selection. In other words, it reverses the work done by a linker. I originally built this as part of a video game decompilation project, having rejected the matching decompilation process used by the community at large. I still needed a way to divide and conquer the problem, which is how I got the funny idea of dividing programs. That allows a particular style of decompilation project I call Ship of Theseus: reimplementing chunks of a program one piece at a time and letting the linker stitch everything back together at every step, until you've replaced all the original binary code with reimplemented source code. It's an exquisitely deep and complex topic, chock-full of ABI tidbits and toolchains shenanigans. There's next to no literature on this and it's antithetical to anything one might learn in CS 101. The technique itself is as powerful as it is esoteric, but I like to think that any reverse-engineer can leverage it with my tooling. In particular, resynthesizing relocations algorithmically is one of those problems subject to the Pareto principle, where getting 80% of them right is reasonably easy but whittling down the last 20% is punishingly hard. Since I refuse to manually annotate them, I've had to relentlessly improve my analyzers until they get every last corner case right. It's by far the most challenging and exacting software engineering problem I've ever tackled, one that suffers no hacks or shortcuts. Once I got it working, I then proceeded in the name of science to commit countless crimes against computer science with it (some of those achievements are documented on my blog). Cross-delinking in particular, that is delinking an artifact to a different platform that it originates from, is particularly mind-bending ; I've had some successes with it, but I sadly currently lack the tooling to bring this to its logical conclusion: Mad Max, but with program bits instead of car parts. Ironically, most of my users are using it for matching decompilation projects: they delink object files from an artifact, then typically launch objdiff and try to create a source file that, when compiled, generates an object file that is equivalent to the one they ripped out of the artifact. I did not expect that to happen at all since I've built this tool to specifically not do this, but I guess when everything's a nail, people will manage to wield anything as a hammer.
- spooneybarger 1y agoThis is insanely cool.
- bloxs 1y ago[dead]
- elorant 1y agoIn my almost 15 years in this community this is the first comment where I don’t have a fucking clue of what is described.
- sureglymop 1y agoWhen you write and compile code, there is a phase called "linking". For example, one could compile pieces of code and then link them together into one binary. A common challenge is decompiling and reverse engineering a compiled program, e.g. a game. The author realized that it's an interesting approach to do a sort of reverse-linking (or I guess unlinking) process of a program into pieces and then focusing on reverse engineering or reimplementing those pieces instead of the whole. When it comes to decompiling games, enthusiasts in that hobby want to be able to reproduce/obtain source code that compiles to exactly the same instructions as the original game. I think there is usually also a differentiation between instruction matching and byte for byte matching reproduction. But it definitely seems like a simpler and better approach to do this piece by piece and be able to test with each new piece of progress. That's my layman understanding of it having only dabbled in decompiling stuff.
- boricj 1y agoThis is pretty much spot on. The key insight of delinking is that object files are relocatable and the process of linking them together (laying their sections in memory, computing all the symbol addresses, applying all the relocations) removes that property. But by reversing this process (creating a symbol table, unapplying the relocations/resynthesizing relocation tables, slicing new sections), code and data can be made relocatable again. Since properly delinked object files are relocatable, the linker can process them and stitch everything back together seamlessly, even if the pieces are moved around or no longer fit where they used to be (this makes for a particularly powerful form of binary patching, as constraints coming from the original program's memory map no longer apply). Alternatively, they can be fed to anything that can process object files, like a disassembler for example. Of course, the real fun begins when you start reusing delinked object files to make new programs. I like to say that it breaks down the linear flow of toolchain from compilation to assembly to linking into a big ball of wibbly-wobbly, bitsy wimey stuff. Especially if you start cross-delinking to a different platform than the original pieces of the program came from.