3 ms·
This was my first thought as well. The question is: How robust are adversarial perturbations? In other words: Given such a perturbation that was generated using
by fantod 6y ago
This was my first thought as well. The question is: How robust are adversarial perturbations? In other words: Given such a perturbation that was generated using one model, how well can we expect to work on a similar model (in the sense that both models are fooled)? I would be curious to know if any research has been done on this question.
- iujjkfjdkkdkf 6y agoMy feeling is that even inception trained on imagenet with a different random initialization would probably not fail (or at least be much less likely to fail) when exposed to these adversarial examples. If these are hacking the specific, underspecified, realization of the classifier (I.e. the set of public weights trained from a specific seed and visit order through the data), the adversarial examples are probably just as fragile as the classifier.
- Filligree 6y agoWhat makes me think is that, in principle, there should be similar adversarial examples for the human optical system. Practically speaking it wouldn't fool anyone for more than a split second, not least since our input is video instead of snapshots, but it's an interesting thing to wonder about. Maybe we could build an AI which would be in most senses as smart as us, but which would be more vulnerable to such things?
- MauranKilom 6y agoProminent examples would be optical illusions and the kind of thing you find on https://reddit.com/r/confusing_perspective/ https://reddit.com/r/confusing_perspective/. Both of which tend to fool for longer than a split-second. It's not entirely the same method as these "adversarial noise" inputs, but some optical illusions are pretty close in how they mess with the localized parts of our optical processing (e.g. https://upload.wikimedia.org/wikipedia/commons/d/d2/Caf%C3%A9_wall.svg https://upload.wikimedia.org/wikipedia/commons/d/d2/Caf%C3%A...). We can't backprop the human vision system to find "nearby" misclassifications as easily, and presumably our own "classifiers" are more robust to such pixel-scale perturbations, but especially lower-resolution images can trip us up quite easily too (see e.g. https://reddit.com/r/misleadingthumbnails/ https://reddit.com/r/misleadingthumbnails/).
- chrisfosterelli 6y agoNot sure on published research, but anecdotally I've played with this and found it depends on the technique used. The most simple attacks tend to be model specific. This often means that it won't work at all on another model or will work but less effectively (the confidence of the adversarial target will increase and the true class will decrease but not necessarily to a degree that will change the classification). It can also depend on the dataset in addition to the model architecture. The more simple ones don't even really work after basic transformations (like rotating, scaling, etc) on the target model, so those attacks are often brittle. But there are lots of techniques and some of them are more robust across more models and transformations. This sometimes has a tradeoff of causing the manipulation to the image to be more noticeable to the human eye. Adversarial attacks are a bit of a cat-and-mouse game between new attacks and new attempts to find where they fail.
- Imnimo 6y agoTransfer of adversarial perturbations between models is one of the main avenues of "black box" (where you don't know the target model's weights) attacks. Perturbations don't translate between models 100% of the time or anything, but many attacks are surprisingly transferable. There are also methods to make perturbations more transferable, for example by finding an attack that is effective against an ensemble of models, you increase the chances that it will transfer to an unseen model. https://arxiv.org/pdf/1611.02770.pdf https://arxiv.org/pdf/1611.02770.pdf
- dheera 6y agoA lot of adversarial attacks tend to correspond to very narrow peaks though and are not robust against some very simple image transformations. Often by slightly disturbing the input image with e.g. blurring or brightness/contrast changes and seeing how the output layer changes you can often eliminate many adversarial attacks. A simple example would be if an image identifies as a basketball but you blur it slightly and it identifies as a cat, you might be looking at an adversarial attack.