4 ms·
These arguments were introduced by Szegedy et. al. earlier this year in this paper: http://cs.nyu.edu/~zaremba/docs/understanding.pdf http://cs.nyu.edu/~zaremba
by zackchase 12y ago
These arguments were introduced by Szegedy et. al. earlier this year in this paper: http://cs.nyu.edu/~zaremba/docs/understanding.pdf http://cs.nyu.edu/~zaremba/docs/understanding.pdf. Geoff Hinton addressed this matter in his Reddit AMA last month.
The results are not specific to neural networks (similar techniques could be used to fool logistic regression). The problem is that ultimately a trained network relies heavily on certain activation pathways which can be precisely targeted (given full knowledge of the network) to fool networks into misclassification on data points which might to a human seem imperceptibly changed from those which are correctly classified. It is important to understand adversarial cases, but unreasonable to get carried away with sweeping pronouncements about what this does or doesn't about all neural networks, let alone intelligence generally, or the entire enterprise of AI research, as seems to happen after a splashy headline.
- bsbechtel 12y agoI know very little about neural networks, other than at the conceptual level, so I could be very off here, but couldn't an algorithm be defined to look for meta signals surrounding the activation pathways which could be regularly 'audited'? What I mean is, humans have a 6th sense that tells us when things are 'off', and require further examination. This comes from having 5 senses that are highly tuned to the world around us, and work incredibly well together. In some ways, physical hardware has limitations on it's ability to 'sense' the outside world, but in other ways, the ability to analyze and collect hard data is significantly more powerful than what humans can achieve, even at a subconscious level. Do current neural network algorithms not have this kind of failsafe?
- pjc50 12y agoHumans don't have a very good conceptual failsafe against known exploits. Everything from optical illusions to political messaging to closeup magic can get past our perceptual filters. In fact optical illusions are probably the best example of this; even knowing that what you perceive is not correct doesn't change your perception.
- bsbechtel 12y agoThat's a great point about human conceptual failsafes, but also, the failsafe I would be referring to in your optical illusion example would be the fact that you actually know what you perceive is incorrect, receiving that information from another data source - someone told you what you are seeing is incorrect. There is never a 100% foolproof failsafe, humans can still easily be manipulated and exploited, but manipulation doesn't work 100% across the board for humans either. Is there value in leveraging failsafes across a network? (I'm just throwing ideas out here)
- aet 12y agoCan you explain how this might be related to the concept of "overfitting"?
- nabla9 12y agoIt's not specially related to overfitting. It's just general property of abstraction process (overfitting can be thought as over-abstraction). Financial accounting can be thought as system that tries to make it easy to differentiate between irregularities and normal operation from few summary documents. The summary documents like balance sheets are abstractions of the original records and you can normally detect financial problems easily using them. But if you have intelligent adversary within the accounting process that tries to forge the result, you must look into the original records and go trough them to find forgery.
- rifung 12y agoThanks for the explanation! I agree that it doesn't seem to be a huge cause of concern for the general use case, but it does make me question the use of machine learning algorithms for computer security, an idea which I've seen a few companies pursuing nowadays.
- Dn_Ab 12y agoI'd characterize this differently and also, as a lot more interesting than that. understanding.pdf can be viewed as a sort of dual to this paper but they're not covering the same thing. In Szegedy et al., they constructed invisibly perturbed images that resulted in the misclassification of previously correctly classified images. Here, the results of a search were images whose classification have little to no visual similarity to typical members of that class. In a way this is interesting because it's a sort of visualization of what the network views as important in discriminating between different objects. It's also interesting as a display of how alien the learned model's view of the world is. Take optical illusions...optical illusions are remotely similar to this sort of exploit, although the sort of scene modeling we do is a lot more complex than recognition or decomposition. Anyways, illusions exploit cues that result in distorted recognition but not drastically so, unlike the case for these networks. My guess is that this is due to animal vision using a lot more high level cues -- cues that are also useful in a natural setting -- depending on things like size, color, shade, lines, context and so on. Visual systems are also a lot more proactive, filtering out things that don't make sense, fudging color at the edges of vision, smoothing out shades and generally making inferences and deductions about what it should be seeing and how things are "supposed" to be. In fact, a good number of illusions exploit those aspects of vision. In the case of these networks, the cues are incomprehensible, having no natural counterpart, so we see most of them as noise. But sometimes they make a kind of sense, as in the starfish, baseball and sunglasses examples. Based on the observations in the paper, I would guess only a handful activations strongly associated to each feature are responsible for each susceptibility. With animal brains the distortions usually end up in a slightly transformed space, a different scaling or something. It's useful to match a bit overzealously and get something like pareidolia but it also makes sense to have the conflations actually be like something you might run into. The ANNs have no such incentive. Their paper also wonders about whether this is unique to discriminative classifiers. Would a generative classifier, with access to a proper distribution, be so susceptible? That'd be very interesting to see. They also mention some real world consequences, some of which I disagree with. Neural Networks are good at interpolating between examples, so if your training has good coverage over what is to be expected then it'll work very well. And in the era of big data this isn't really a problem (that they don't generalize as we do might explain some of why they have trouble with abstract images) so I'm skeptical an image search solution would be thrown off by textures. There is, however, a better example of facial or speaker recognition. For example, you could train a network to distinguish between faces or voices and then evolve a pattern against it. This could then be used in such a way as to be randomly matched to an individual on a target database. Not good. Driverless cars are also mentioned but those are typically augmented beyond just vision. Personally, I'd add medical scans to the list of things to be careful with. Finally, it's worth mentioning that some of the evolved images are inspired works of art. And a few of the images optimized (not evolved) with an L2 penalization are recognizable without the label and a few more where you can see why it gave the label it did. Your offhanded dismissal was unwarranted IMO.
- Bill_Dimm 12y ago"The results are not specific to neural networks (similar techniques could be used to fool logistic regression)." Do you have any references or examples where such behavior has been demonstrated on non-NN algorithms (like logistic regression)?
- Bill_Dimm 12y agoOK, I found where Hinton claimed that[1], but it would be nice to have a better reference than a Reddit comment if anyone knows of one. [1] http://www.reddit.com/r/MachineLearning/comments/2lmo0l/ama_geoffrey_hinton/clyjbai http://www.reddit.com/r/MachineLearning/comments/2lmo0l/ama_...
- chestervonwinch 12y agoI can verify this is true. It is not extremely difficult to build a multinomial regression model and find "adversarial" images for this trained model. Of course, you're just taking me for my word on this, but you could very well try it yourself. You don't really need any special software outside the tools of SciPy and sklearn for instance if you use python.
- michaelochurch 12y agoQuestion: do these blind spots of neural nets persist as the data set is expanded? I understand that if you have a fixed training data set of, say, 10 million images, there are probably going to be adversarial examples that will be specific to a net trained on that set. (These may be the same artifacts that are picked up by overfitting, or they could be local underfitting artifacts due to early stopping.) But are they specific to the data set? In other words, are the same adversarial examples going to "break" a neural net trained on a separate (but identically generated) set of 10 million images? Or is this something that can be smoothed away with, say, ensemble methods and multiple data sets?
- pfortuny 12y agoI guess this boils down to something similar to the approximation of a continuous function by polynomials: given any finite number of values (x_i, y_i), you can always find a polynomial passing through them. However, there are continuous functions which are nowhere differentiable and which have unbounded variation between x_i and x_{i+1} for each i. The "polynomials" are the "nice images" and the "continuous with unbounded variation" are the "noisy ones". Disclaimer: I am neither a lawyer not an expert on NN. Edit: I would like to be corrected if I the similitude is wrong.
- mathattack 12y agoEntire fields of finance exist for this. Find the hole in the model, and stuff billions of bonds and derivs in it. This is especially an issue when the model sellers optimize off of it. Also a reason why it's good to be leery when people say, "I don't need to know what's going on in the model, the results are good enough."