4 ms·
I've been thinking on how to build a benchmark for this stuff for a while, and don't have a good idea other than LLM-as-judge (which quickly gets messy). I gues
by rfoo 2y ago
I've been thinking on how to build a benchmark for this stuff for a while, and don't have a good idea other than LLM-as-judge (which quickly gets messy). I guess there's a reason why current neural decompilation attempts are all evaluated on "seemingly meaningless" benchmarks like "can it recompile without syntax error" or "functional equivalence of recompilation" etc.
- vessenes 2y agoHmm, specifically when it comes to reverse engineering, you have the best benchmark ever - you can check the original code, no?
- brokensegue 2y agothat requires LLM as judge
- dataangel 2y agono it doesn't, you just diff against the real source code. probably something more fuzzy/continuous than actual diff, but still
- brokensegue 2y agoProving that two pieces of code are equivalent sounds very hard (incomputable)
- rfoo 2y agoBesides functional equivalence, a significant part of the value in neural decompilation is the symbol (function names, variable names, struct definition including member names) it recovered. So, if the LLM predicted "FindFirstFitContainer" for a function originally called "find_pool", is this correct? Wrong? 26.333% correct?
- bitfieldz 2y ago[dead]