3 ms·
The most difficult parts of getting readable code would be dealing with inlined functions and otherwise-duplicated code from macros or similar, and dealing with
by dzaima 11mo ago
The most difficult parts of getting readable code would be dealing with inlined functions and otherwise-duplicated code from macros or similar, and dealing with in-memory structure layouts; both pretty complicated very-global tasks. (never mind naming things, but perhaps LLMs have a good shot at that)
That said, chatgpt currently seems to fail even basic things - completely missed the `thrM` path being possible here: https://chatgpt.com/share/69296a8e-d620-800b-8c25-15f4260c78db https://chatgpt.com/share/69296a8e-d620-800b-8c25-15f4260c78... https://dzaima.github.io/paste/#0jZJNTsMwEIX3OcWoSFWCqrhN0wbKArkgVlAJAevKf2ktJU5lO2oK6jEQK/ZwDi7EEXCUNoDogs3I9rz3@XlkhEBUK6KMLBQUKSytXZkJQgtplyUNWZEj/khkTtDF9HaGaFZQNBrzPh/SIYkFoVE/PklIwlI6GDF6GpE@SXjimsMEGc0QLWVmpTKIaIfkITu6jgeRZyyxkoFUmVQCppCX2fwO@1OwPbdb16UK4MkD0MKWWgHf4BpQi7pOfLm5JzQTM9yrxVVw5m0PMvF/mPgXE/9keg2GRYcwMgVfmqtx7K8D6EKzrIKm2d6Sz9NaEaZwDFWYOizAtnVjrf9ammMHPd8/wu4ywWQ/KtumdDgQmRE7xjd37bj7EK2pcr42g13qG7/z@fr89vHi6vsEHpT7D4JZwYHoRZkLZcFuVsJ06nl8AQ#C https://dzaima.github.io/paste/#0jZJNTsMwEIX3OcWoSFWCqrhN0wb... and that's only basic bog-standard branching, no in-memory structures or stack usage (such trivial problems could be handled by using an actual proper disassembler before throwing an LLM at that wall, but of course that only solves the easy part)
- CamperBob2 11mo agoYou can't feed something like that to the free ChatGPT model and expect anything useful. Try these: https://chatgpt.com/s/t_6929f00ff5508191b75f31e219609a35 https://chatgpt.com/s/t_6929f00ff5508191b75f31e219609a35 (5.1 Pro Thinking) https://claude.ai/share/7d9caa25-14f7-4233-b15c-d32b86e20e09 https://claude.ai/share/7d9caa25-14f7-4233-b15c-d32b86e20e09 (Opus 4.5) https://docs.google.com/document/d/1C0lSKbLSZOyMWnGgR0QhZh3QxwdIZmSufNr6H5sXwQg/edit?usp=sharing https://docs.google.com/document/d/1C0lSKbLSZOyMWnGgR0QhZh3Q... (Gemini 3 Pro Thinking) All of them recognized the thrM exception path, although I didn't review them for correctness. That being said, I imagine the major showstopper in real-world disassembly tasks would simply be the limited context size. As you suggest, a standard LLM isn't really the best tool for the job, at least not without assistance to split up the task logically.
- dzaima 11mo agoThose first two indeed look correct (third link is not public); indeed free chatgpt is understandably not the best, but I did give it basically the smallest function in my codebase that does something meaningful, instead of any of the actually-non-trivial multi-kilobyte functions doing realistic things needing context.
- CamperBob2 11mo agoWould be interesting to push the models with a couple of larger functions, if you have some links you'd like me to try. I have paid pro accounts on all three, but for some reason Gemini is no longer allowing links to be shared on some queries including this one. All it would let me do is export it to Docs, which I thought would be publicly visible but evidently isn't.
- dzaima 10mo agoActually, even finding a larger function that would by itself have a meaningful disassembly is posing problematic; basically every function deals with in-memory data structures non-trivially, and a bunch do indirect jumps (function pointers, but also lookup-table-based switches, which require table data from memory in addition to assembly to disassemble). Like, here's a ~2.7x larger function: https://dzaima.github.io/paste/#0jVdNjxs3DL3nVwzQo30gRY00ChYLpEjPTdKgRREsgvlEvF3vGra3cPPrK1IzNmc8uzs6GRyKIvken@SbFdy@z@rtLourPWzWcMJ3NyuKxvuWbRmcrIfOQ3bTbnfH//5s69voEKLD7vnwIzrs0UYDYrScTdWJTUabSjFx4Gz79C9HLh/iaQRsztlc/fyxSZ7rfbNZpx0G@h0/TzHB@rT@9e@vv2Wfvn7JvkWHFZzcHXtZ7dWMvGr2yssOcmBXkkSHiqNr/MhmiXBftkPNCI3pa/6w368Kw2WTU/lzop//@v3Lx/6gzS4ehNbEnXdZv36JsfKKsCqym@32e/Vc/9MeDysLHM1KN47t4TiULSVbOUS1H7HpVCqUI@/Ojc6lHucirSm43lyibR5rdptLlyBM06U6pVsf90OuTiM3FwY6c7eOaSyq3A@oltVBco8hWoAIUVrMwUJ87hNQqQ3gVBswQsZu0oayaS54dmlV/DW/QpvYXEzRRgw6dk7CclgANzR@GdzBz8KNAONpQ2wrDbeTXGLx0e1NuGOTotsrcJuGFsCN6C@nzVZd1cvhRpnPMdzlBG4kO39ieZKDxMW/5CLlx7CiJpZB7xmxP@yiuRCziNaTUKFXqByUKemYjNW@PcpvUmlXnLbvOgOXxcqBMhoDxaS4RsxuCiuN6JsUBV2hYEW77jPzOA2KTuzcpftHRdyiVFFNPFXc3DjqVbcwsUWGbDRjVI6STOFkyF4LZ/pwTjWsVLOYFrsEp8DhIUhtD2HKED/ZaSBcWiI7hRRGrpnRKDvUHTH5rbgpel1uHoOcjr55GIJ08xi8TNLHMSHlM01b50ZIJHni49@fh56DS73GFBN2QGlHQy@byc71UxFQarDsVT5OWmqsn2kpTjfnfrarbipK0Ojqgsy1EY4Pm8/kNY6ROrTHVnrNFq8IVD8MbDbC5sNz1Zt7X83dl0QXF96xpiDV/7PommI6nQZ0fRQSeoWezuYl0TXBvCG61@nOiK4ZhuBF0QUXtbBZJroEehjTQ0fpLaF7RW@bk7iofK7fXeua4SJjlogyEcxRhYimKol6elnLxS28ckjUsp7zJE@qK@knGYWR9JMtptJPchsk6ec31vnAttkMr8RcNa1/LrfyLsSpGHS5ppOn9Hyk2aj6Duijlmz3crOUDw8pqivbDgTpaHqq/5CQvljU/sLNNibAVWMCXjUm0KUxoZckeWnHXBMJgp8rzIJ5swALiy51K/8RrvhjsVjCHysPkBE@DWl8Ir/e/Q8#asm https://dzaima.github.io/paste/#0jVdNjxs3DL3nVwzQo30gRY00ChY... (is https://github.com/dzaima/CBQN/blob/90c1dc09e88c5324373281f6c5d76811ca4439f3/src/core/fillarr.c#L218-L227 https://github.com/dzaima/CBQN/blob/90c1dc09e88c5324373281f6... with a bunch of inlining) (I'm keeping the other symbol names there even though they'd likely not be there for real closed-source things, under the assumption that for a full thing you'd have something doing a quick naming pass beforehand) This is still very much on the trivial end, but it's already dealing with in-memory structures, three inlined memory allocation calls (two half-deduplicated into one by the compiler, and the compiler initializing a bunch of the objects' fields in one store), and a bunch of inlined tagged object manipulations; should definitely be possible to get some disassembly from that, but figuring out the useful abstractions that make it readable without pain would probably take aggregating over multiple functions. (unrelated notes of your previous results - claude indeed guessed correctly that it's BQN! though CBQN is presumably wholesale in its training data anyway; it did miss that the function has an unused 0th arg (a "this" pointer), which'd cause problems as the function is stored & used as a generic function pointer (this'd probably be easily resolved when attempting to integrate it in a wider disassembly though); neither claude nor cgpt unified the `x>>48==0xfff7` and `(x&0xffff000000000000)==0xfff7000000000000` which do the exact same thing but clang is stupid [https://github.com/llvm/llvm-project/issues/62145 https://github.com/llvm/llvm-project/issues/62145] and generates different things; and of course a big question is how many such intricacies could be automatically reduced down with a full codebases worth of context, cause understandably the single-function disassemblies are way way more verbose than the original)