3 ms·
Authorship style identification in natural language has good intuition that one can work with. However, such notion in binary executive sounds totally nonsense.
by LambdaTrain 6y ago
Authorship style identification in natural language has good intuition that one can work with. However, such notion in binary executive sounds totally nonsense. The only possibility that makes it work might originate from code reuse in a single organization, which is a classical feature to look into for malware detection.
So I have no idea what new information this arxiv paper provides other than to introduce an academic topic full of fancy terminologies.
- vardump 6y agoI wouldn't be so quick to dismiss this possibility. Sure code compilation (and LTO) is going to filter out a lot of the signal, but you still have potential unique patterns in the structure of the system and in the data flow. At least it can still be used together with other methods to compute a probability between different authors and organizations.
- jstanley 6y agoI'd say code compilation is a lot of the signal, in the sense that some people's build environment is likely to be quite unique. It's like identifying people based on their browser fingerprint.
- vardump 6y agoA good point. Definitely any build environment fingerprints are useful data, like compiler and linker quirks. What I should have said was that a different compiler (and/or version) can filter the original input in different ways, because of optimizations and other code transforms (including non-breaking codegen bugs).