5 ms·
It is good to know that they need access to a lot of predictions from a net, before they can create an image that will "fool" the net, but look alien to humans.
by compbio 12y ago
It is good to know that they need access to a lot of predictions from a net, before they can create an image that will "fool" the net, but look alien to humans. Secondly, this doesn't account for ensembling: "fool me once, shame on you. Fool me twice...". Since the images are crafted for a single net, a majority vote should not be fooled by these images. I suspect this effect rapidly goes away when adding more nets (which is basically industry-standard practice to increase accuracy).
Furthermore, I am seeing the security concerns, but I figure this is far from a practical attack. Deep Learning Classifiers do not act as gatekeepers: You have not much to gain from a single faulty classification. You won't be granted access to secret information if you happen to look like the CEO.
- userbinator 12y agoFurthermore, I am seeing the security concerns, but I figure this is far from a practical attack Perhaps it's a sign of the times that almost every discovery that could possibly be related to security in some way, does. I have a feeling that if this was a decade or two ago, the sentiment would be very different. ("Can you figure out what a computer thinks these images are?") Also, the image labeled "baseball" immediately reminded me of a baseball...
- deleted 12y ago[deleted]
- pavel_lishin 12y ago> Since the images are crafted for a single net, a majority vote should not be fooled by these images. I suspect this effect rapidly goes away when adding more nets Sure, but this assumes that whatever neural net system you're relying on was bought by someone who is more security conscious than they are cheap.
- jerf 12y ago"Secondly, this doesn't account for ensembling: "fool me once, shame on you. Fool me twice...". Since the images are crafted for a single net, a majority vote should not be fooled by these images." Trivially "solved" by treating the ensemble as a single object, then constructing a counterexample. My intuition suggests that while the resulting "fooled you" image may very slowly converge on something human recoginizable, it won't do so at a computationally-useful rate.
- compbio 12y agoNow I have to try this out. My intuition tells me it becomes increasingly hard to create a fooling image, which looks alien, and is able to fool all the nets in the ensemble, even though they have different settings and params. I think they can only fool one net at a time, and have to get very lucky to be able to evolve the image for the other nets, while keeping the same classification. You can't "train" these images on all nets at once, by simply treating the ensemble output as a single net. If your intuition is right though, then the ensemble may be able to counter with a random selection of nets for its vote: You'd need to evolve images for every possible combination and/or account for nets added in the future.
- Houshalter 12y agoOverfitting has nothing to do with it. See this paper: http://arxiv.org/abs/1412.6572 http://arxiv.org/abs/1412.6572 I believe the original paper tried ensembles and even got the images to work on different networks.
- compbio 12y agoI was not talking about overfitting. I've seen that paper. The original paper asked if images that could fool DBN.a could fool DBN.b. The answer was: certainly not all the time. They used the exact same train set and architecture for DBN.a and DBN.b, just randomly varied initial weights. I think this is too favorable for a comparison with a voting ensemble made with nets with a different architecture, train set and tuning. Can they also find images that can fool DBN.a-z? Also, to test if a net can learn to recognize these fooling images, they simply add them to the train sets. Those noisy images would be far simpler to detect: They have a much greater complexity than natural images. To detect the artsy images, a quick knearest-neighbors run should show that they do not look much like anything it has seen before, so it may be an adversarial image.
- Houshalter 12y agoTo be clear I meant this paper (http://arxiv.org/abs/1312.6199 http://arxiv.org/abs/1312.6199) as the original paper for adversarial images. I think they did try transferring them between very different NNs: >In addition, the specific nature of these perturbations is not a random artifact of learning: the same perturbation can cause a different network, that was trained on a different subset of the dataset, to misclassify the same input.
- scribu 12y ago> You won't be granted access to secret information if you happen to look like the CEO. Tell that to this startup: http://onevisage.com/ovi-technology-page/ http://onevisage.com/ovi-technology-page/
- jeffclune 12y agoSurprisingly, ensembles do not really help. We have tried this and it does not work (the final paper for CVPR 2015 will show these results). Also see the work of Szegedy et al. and Goodfellow et all, which also show that ensembles do not really help.
- jeffclune 12y agoAlso, you don't necessarily need a lot of predictions from a net to fool it, because (another surprising result) images that fool one net tend to fool others! So I can create fooling images on my in-house net and then take them and fool your net-used-for-some-important-application without getting any feedback from your net. That's very surprising, and does raise serious security concerns.