5 ms·
This article left me thinking about the possibility of adversarially constructed malware. It seems like an adversarial network could modify existing malware to
by rectangletangle 9y ago
This article left me thinking about the possibility of adversarially constructed malware. It seems like an adversarial network could modify existing malware to look like benign machine code from the classifier's perspective. This might result in an "arms race," similar to text spinners vs spam classifiers.
Regardless this seems like a promising technique, even with that potential caveat. Since most malware out in the wild isn't that sophisticated, this is likely quite effective.
- discreditable 9y agoAnother interesting possibility could be an adversary training the network to treat benign code as malicious.
- rectangletangle 9y ago"Why does the pop up keep saying it's a virus?" "I don't know, it always says that. Just click 'proceed anyway'"
- abcd_f 9y agoThat's not a "possibility", it's a reality. From what we are seeing (as a desktop software vendor) all fancy-shmancy AI-based antiviruses absolutely "excel" at false positive detection. It's more of a miracle when they do NOT flag something that's not of "hello world" variety as a malware. And I wish I were kidding.
- LeoJiWoo 9y agoThat is concerning. I knew over-fitting was a real issue with AI. Still haven't found a good way to deal with high bias or high variance myself.
- rectangletangle 9y agoCross validating the classifier/hyper parameters and a good scoring metric (Matthews correlation coefficient) go a long way. Since the classes are very imbalanced, an appropriate scoring metric is very important. Even more importantly, train with lots of high-quality data whenever possible. Anecdotally many seem to obsess over the particular classification algorithm, while neglecting data quality. A classifier is only ever as good as its training set.
- EdwardRaff 9y agoPaper author here! A lot of that issue comes from people using bad datasets. One of our first papers was about that ( http://www.readcube.com/articles/10.1007/s11416-016-0283-1?author_access_token=Y2ftVow3BBIXRTHYIxoCG_e4RwlQNchNByi7wbcMAY4NW74db1mhZZQDQYJ1tM7Y-KZqnwIXRhZC64F6SuX0bowkkoy4Ro-NFZSGOs2sw2kG7I6cMZb9G3I0tfGpLO_rZlh-MF7KZ2i-qxjmAi-Shw%3D%3D http://www.readcube.com/articles/10.1007/s11416-016-0283-1?a... ), and showed that using the data most people use in their research, benign data collected from clean Microsoft installs, is not sufficient. The model will literally learn to look for the string "Copyright Microsoft Corporation" to decide if something is benign. Everything else ends up getting marked as malicious. We are using better data in this work, and it does not suffer from this problem. It is not ready to be a real production AV, but it does a fairly good job at separating out benign vs malicious files and dealing with non-trivial examples of both.
- zitterbewegung 9y agoWhat about malware that when you scan it it will trigger an RCE by exploiting the classifier inside of the virus scanner ?
- EdwardRaff 9y agoThis is something we are looking at! It is a harder problem to create adversarial examples in the malware space, because you can't make arbitrary changes and have the code still work. Endgame has a great paper on this problem, and showed how they can defeat regular AVs with some fairly simple modifications that don't impact the malware's execution. https://www.blackhat.com/docs/us-17/thursday/us-17-Anderson-Bot-Vs-Bot-Evading-Machine-Learning-Malware-Detection-wp.pdf https://www.blackhat.com/docs/us-17/thursday/us-17-Anderson-...