9 ms·
Deep Neural Networks Are Easily Fooled
- akiselev 12y agoWe humans are as brilliant at pattern matching as we are in finding patterns that aren't really there, not just with our vision but with our understanding of probability , randomness, and even cause & effect. Thankfully, our brains are very complicated machines that can recognize a stucko wall or a cloud and invalidate the false identification of a face or unicorn or whatever based on that context. With that in mind, is it really surprising that [m]any of our attempts at emulating intelligence can be easily fooled? An untold number of species have evolved to do exactly the same thing: exploit the pattern matching errors of predators to disguise themselves as leaves or tree branches or venomous animals that the predator avoids like the plague. DNNs seem to be relatively new and we've got a long ways to go, so is this a fundamental problem with the theoretical underpinning or do we just need to train them with far more contextualized data (for lack of a better phrase)? Is there any chance of us having accurate DNNs if we can, as if gods during the course of natural selection, peek into the brain of predators (algorithms) and reverse engineer failures (disguises for prey) like this?
- Dewie 12y agoWhat I don't get about AI optimism: So we humans are very fallible and can easily make mistakes. Computers are better at us when it comes to problems that are clearly defined and for the problem is decidable (can implement an algorithm to solve any instance of the problem). But we need AI when the problems are more fuzzy, like recognizing a lion in a picture. How can we build a mostly automated future, if the AIs that are supposed to do our jobs turn out to be very fallible as well? They won't - supposedly - have the problem of being self-aware and being able to follow their emotions rather than their own best judgement and reasoning. But it seems that some problems are inherently prone to making mistakes. Can it be avoided at all? And if so, who do we blame when an AI makes a "mistake" like that? The training set?
- Alphasite_ 12y agoI'd think all you can do is design for failure (as you would now) and use the failures as training data. The only real advantage is that you don't pay AI and they don't get bored, they're consistsntly good or bad.
- derefr 12y agoPeople's fallibility goes down as you throw more of them at a task—not because the majority will be right, but because the signal adds up while the noise cancels out. This is what the efficient market hypothesis, "wisdom of crowds", etc. are basically about. If you train 1000 AIs on different subsets of your training corpus, their ensemble will be much "hardier" than one AI trained on the entire corpus. The automated future comes from the fact that you didn't need 1000 full training corpii to get this effect, nor do 1000 AIs cost much more than one to run, once you've built out hardware enough for one. In other words, AI makes the application of "brute-force intelligence" to a large problem cheap enough to be feasible, in the same way slave labor made building pyramids by brute force cheap enough to be feasible.
- slavak 12y agoExcept, of course, that the pyramids are a marvel of workmanship and engineering probably built by a relatively small force of skilled workers. Their construction methods are about as far from "brute-force" as you could possibly get. http://science.howstuffworks.com/engineering/structural/pyramid2.htm http://science.howstuffworks.com/engineering/structural/pyra...
- nightski 12y agoThere are many examples where crowd behavior exhibits less "wisdom". Take any market bubble for example. Have we ever looked at a tough scientific problem and came to the conclusion that the best path forward was to collect as many random people off the street and shove them in a room to solve it? Also bootstrapping or model parameter selection techniques are already heavily used in AI and have not yet brought us this future. I believe that the model you presented has been simplified a bit too much ignoring a lot of important variables.
- Animats 12y agoThere was a similar result a few months ago for another type of machine learning. (That's note 26 in this paper.) The problem seemed to be that the training process produces results which are too near boundaries in some dimension, and are thus very sensitive to small changes. Such models are subject to a sort of "fuzzing attack", where the input is changed slightly and the output changes drastically. There are two parts of this process that are kind of flaky. The problem above is one of them. The other part is feature extraction where the feature set is learned from the training set. The features thus selected are chosen somewhat randomly and are very dependent on the training set. It's amazing to me that works at all. Earlier thinking was to have some canonical set of features (vertical lines, horizontal lines, various kinds of curves, etc.), the idea being to mimic early vision, the processing that happens in the retina. Automatic feature choice apparently outperforms that, but may not really be working as well as previously believed. It's great seeing all this progress being made.
- simonster 12y agoAutomatic feature choice can actually lead to a set of features that resembles V1 receptive fields (as demonstrated by Olshausen and Field in https://courses.cs.washington.edu/courses/cse528/11sp/Olshausen-nature-paper.pdf https://courses.cs.washington.edu/courses/cse528/11sp/Olshau...). I recently attended a talk by Geoff Hinton on "capsules." He pointed out that the max pooling used in convolutional neural networks effectively disregards information about relationships among features. Instead, he propose a network composed of "capsules" that each estimate whether an implicitly defined intermediate feature is present and its pose. The idea is that an object is present only if its intermediate features are present and their poses agree. He showed some neat results from these models (some published in http://arxiv.org/pdf/1412.1897v1.pdf http://arxiv.org/pdf/1412.1897v1.pdf, and some from http://www.cs.utoronto.ca/~tijmen/tijmen_thesis.pdf http://www.cs.utoronto.ca/~tijmen/tijmen_thesis.pdf). Notably, these models can evidently learn to classify MNIST with >98% accuracy given only 25 labeled examples. (I am not sure how many unlabeled examples were used.) I don't have any experience with these models, but given that most of these images look like a single feature embedded in noise or as a texture, I would not be surprised if a capsule-based network would not be so susceptible to these images.
- cLeEOGPw 12y agoThe vulnerability exploits imperfections in the NN weights. To avoid this kind of mismatch all you need to do is shift the same image by 1 pixel (assuming recognition is done per pixel), and you can cross check results to check if an error occurred. Human brain recognizes better because it can sample the image many times from many slightly different angles. There's a reason saccade (http://en.wikipedia.org/wiki/Saccade http://en.wikipedia.org/wiki/Saccade) exists.
- phreeza 12y agoThis is highly oversimplified. The previous example of antagonistic samples being found works even if the training data was perturbed as you described. The reasons for saccades are also a lot more complex, involving photoreceptor adaptation etc.
- alinenache 12y agohttp://waltherpragerandphilosophy1.blogspot.ro/2013/01/improving-our-species_11.html http://waltherpragerandphilosophy1.blogspot.ro/2013/01/impro...
- bitL 12y agoSo, deep neural networks are like artists, able to see a structure in chaos? Like when Michelangelo looked at a large stone, seeing David there immediately, so do DNNs recognize lions in white noise? We should applaud introduction of phantasy and imagination into science ;-)
- zackchase 12y agoThese arguments were introduced by Szegedy et. al. earlier this year in this paper: http://cs.nyu.edu/~zaremba/docs/understanding.pdf http://cs.nyu.edu/~zaremba/docs/understanding.pdf. Geoff Hinton addressed this matter in his Reddit AMA last month. The results are not specific to neural networks (similar techniques could be used to fool logistic regression). The problem is that ultimately a trained network relies heavily on certain activation pathways which can be precisely targeted (given full knowledge of the network) to fool networks into misclassification on data points which might to a human seem imperceptibly changed from those which are correctly classified. It is important to understand adversarial cases, but unreasonable to get carried away with sweeping pronouncements about what this does or doesn't about all neural networks, let alone intelligence generally, or the entire enterprise of AI research, as seems to happen after a splashy headline.
- bsbechtel 12y agoI know very little about neural networks, other than at the conceptual level, so I could be very off here, but couldn't an algorithm be defined to look for meta signals surrounding the activation pathways which could be regularly 'audited'? What I mean is, humans have a 6th sense that tells us when things are 'off', and require further examination. This comes from having 5 senses that are highly tuned to the world around us, and work incredibly well together. In some ways, physical hardware has limitations on it's ability to 'sense' the outside world, but in other ways, the ability to analyze and collect hard data is significantly more powerful than what humans can achieve, even at a subconscious level. Do current neural network algorithms not have this kind of failsafe?
- pjc50 12y agoHumans don't have a very good conceptual failsafe against known exploits. Everything from optical illusions to political messaging to closeup magic can get past our perceptual filters. In fact optical illusions are probably the best example of this; even knowing that what you perceive is not correct doesn't change your perception.
- 12y ago
- benanne 12y agoThe discussion about this paper on r/MachineLearning is quite insightful and worth reading: http://www.reddit.com/r/MachineLearning/comments/2onzmd/deep_neural_networks_are_easily_fooled_high/ http://www.reddit.com/r/MachineLearning/comments/2onzmd/deep...
- ifdefdebug 12y agoI know literally nothing about this science, so that paper had me concerned about the following question: Given a visual face recognition door lock or similar system. If I want to break such a door lock, can I install that system at home, train it with secretly taken pictures of an authorized person, and evolve some kind of key picture with my home system until I can show it to the target door lock and fool it into giving me access? OK this is a very simplified way to put the question, but is that something this paper would imply to be possible (in a more sophisticated way)?
- RyanZAG 12y agoYes, exactly. However, you will need the database or training data fed into that lock system. Without the data, you won't be able to test it at home.
- TeMPOraL 12y agoYou could probably arrange that person to be on a photograph (say ask him/her to hold a sign "I support our war heroes!" or something else everyone would agree to hold if you said you're making a photo album for a non-profit) and work from there. But in general, this class of problems is why a biometric lock is only useful if accompanied by a guard with a gun stationed next to it.
- compbio 12y agoYou may be interested in this talk: https://www.youtube.com/watch?v=tleeC-KlsKA#t=282 https://www.youtube.com/watch?v=tleeC-KlsKA#t=282 The speakers runs you through a hypothetical case study: a pet-door company and looks at the pitfalls of applying machine learning to it. I believe the paper is actually focusing on something else: Create images that humans will not be able to classify as a digit, but that the net will gladly give a prediction for. To translate to faces: There may be clouds that look random to our human brain, but are detected as faces by nets. It seems this is an adversarial attack, where you need access to the guts of the net (weights, layers). I compare this with a hashing algorithm and brute-forcing the input till you find a collision with a target. Nearly impossible in real-life situations. You may be able to sign a check using a scribble that the cashier can not recognize as a digit, but that the machine will recognize as a digit. Not much practical gain from an attack there.
- robg 12y agoSo are human brains.
- yummyfajitas 12y agoI wish they explained why evolutionary algorithms were used. They seem to suggest gradient ascent also works - I wonder what the key criteria are for constructing good adversarial images?
- deleted 12y ago[deleted]
- hippich 12y agoThis brings interesting question. Is it possible to hack human brain? Will specific set of stimuli make brain react in certain way?
- compbio 12y agoA field in cognitive science that asks these questions is "gestalt psychology". http://en.wikipedia.org/wiki/Gestalt_psychology http://en.wikipedia.org/wiki/Gestalt_psychology
- monochr 12y agoI have nothing intelligent to say without reading the full paper... ...But, how different is this from the various optical illusions humans fall for? I mean we can't exactly tell the difference between a rabbit and duck ourselves[1] so isn't it just a universal property of all neural-network like systems that there will be huge areas of mis-classifications for which there hasn't been specific selection? [1] http://mathworld.wolfram.com/Rabbit-DuckIllusion.html http://mathworld.wolfram.com/Rabbit-DuckIllusion.html
- praptak 12y agoI believe that the point of the article is that the triggers for optical illusions are totally different in humans and the ANNs. I don't know how valid is this statement - humans sometimes do recognize "shapes" in white noise too.
- deleted 12y ago[deleted]
- larrydag 12y agoAnother journal paper covering the same thing. http://arxiv.org/abs/1312.6199 http://arxiv.org/abs/1312.6199. And the article I got that references it. http://www.i-programmer.info/news/105-artificial-intelligence/7352-the-flaw-lurking-in-every-deep-neural-net.html http://www.i-programmer.info/news/105-artificial-intelligenc...
- krick 12y agohttps://news.ycombinator.com/item?id=8544911 https://news.ycombinator.com/item?id=8544911
- SeanDav 12y agoIf the shoe was on the other foot, I can imagine a race of super computer AI's administrating a similar test to humans and saying look at the puny human vision system. It is fooled easily by simple optical illusions that wouldn't fool even a 2 year old AI. Clearly, there are questions about the generality of the human vision system and perhaps it is not fit for purpose...
- romaniv 12y agoYes, of course, what constitutes sensible image recognition is just a matter of opinion, and opinions of some algorithm are as valid as yours. In fact, my pseudo-random number generator seeded with image binary accurately recognizes 100% of images I give it (based on my newly developed definition of image recognition).
- SeanDav 12y agoI have got no idea of the point you are trying to make. You seem to be criticizing something I said but not sure what. You do realize that I was trying to make a somewhat ironic / somewhat humorous comment that you have to be careful that what you test for is relevant and that one isn't focusing on a narrow weakness which may not actually be that relevant for the general case.
- romaniv 12y agoYou're making a comment in response to a specific research paper. Therefore, I interpret the comment within the context of that paper. So you're implying that the paper is "focusing on a narrow weakness which may not actually be that relevant for the general case". I disagree. Any 2d image is an optical illusion, so it makes no sense to criticize human image recognition based on it being 'fooled' by illusions. The real criteria for whether image recognition works well or not is altogether different.
- crimsonalucard 12y agoIf we could find out the selection criteria behind each layer of the neural network for the human visual cortex we could possibly build something more accurate. Although I doubt the visual cortex is a simple feed forward network like the one used in the paper. It's likely to have a non linear structure that's significantly more complex.
- jacobsimon 12y agoI don't have too much of a problem with this actually, because a lot of the "nonsense" images actually bear strong resemblance to the objects. The gorilla images clearly look like a gorilla, the windsor tie images clearly show a collar and a tie. The image coloring is way off of course, but the gradients seem about right.
- comex 12y agoOne might say that Picasso's Bull is a human equivalent of this: he "evolved" a sequence of images and ended up with something that has very few features of a bull, but nevertheless gets recognized by humans as such. Then again, unlike the neural networks in the paper, humans would be capable of classifying abstract images into a separate category if asked.
- MrQuincle 12y agoI start to like this way of looking for false positives or false negatives more and more. It would be interesting to introduce some kind of aspects known from the human brain and see if the misclassified items "move" in some conceptually understandable direction. * Introduce time. Humans are not just image classifiers; humans are able to recognize objects in visual streams of images. Such streams can be seen as latent variables that introduce correlations over time as well as space. What constitutes spatial noise might very well be influenced in our brains by the temporal correlations we see as well. * Introduce saccades. A computer is only able to see a picture from one viewpoint. Our eyes undergo saccades and microsaccades. That's an unfair advantage for us, being able to see a picture multiple times from different directions! * Introduce the body. We can move around an object. This again introduces correlations that 1.) are available to us, and 2.) might define priors even when we are not able to move around the picture. In other words, we can (unconsciously) rotate things in our head.
- sxyuan 12y agoYou might be interested in this paper, if you haven't seen it already: http://arxiv.org/abs/1406.6247 http://arxiv.org/abs/1406.6247
- fallenpegasus 12y agoWhat this tells me that there probably exist deeply weird images that would be recognized as something by one person or by very few people, but would be just an unrecognizable mash of colors and lines to everyone else.