7 ms·
Why do CNNs generalize so poorly to small image transformations?
- PredictorY 8y agoIt's worth noting that the actual title of that essay is, "Why do deep convolutional networks generalize so poorly to small image transformations?"
- sctb 8y agoThanks! We've un-generalized the submission title from “Why do neural networks generalize so poorly?”.
- denzil_correa 8y agoOne of the cited papers "Measuring the tendency of CNNs to Learn Surface Statistical Regularities" is also a great insight into this phenomena [0]. > Our main finding is that CNNs exhibit a tendency to latch onto the Fourier image statistics of the training dataset, sometimes exhibiting up to a 28% generalization gap across the various test sets. Moreover, we observe that significantly increasing the depth of a network has a very marginal impact on closing the aforementioned generalization gap. Thus we provide quantitative evidence supporting the hypothesis that deep CNNs tend to learn surface statistical regularities in the dataset rather than higher-level abstract concepts. [0] https://arxiv.org/abs/1711.11561 https://arxiv.org/abs/1711.11561
- calebh 8y agoI've recently become interested in permutation invariant neural networks. There has been very little work in this area - just PointNet and a few derivatives. Anyway, I think that neural networks are now entering the trough of disillusionment as people begin to discover the limitations. Maybe in the future, somebody will come up with a new machine learning architecture that has better generalization. I'm not expecting gradient descent to give us general AI.
- salawat 8y agoThe main cause of the generic brittleness of Neural Networks is probably in the way they are utilized. Biological neural nets never really stop learning. They slow down, even "forget" in order to restructure, but they change constantly. A static neural net in basically a snapshot of it's environment (training data). Very interesting consequences for the ML field if my hunch has anything remotely resembling a kernel of truth to it.
- sova 8y agoContinue learning while also continuing to "forget" ... that's very interesting! Space required to save data remains constant, and their ability adapts to environment and conditions. So... seeing neural nets as living structures instead of as snapshot sieves/filters... very strong approach. I wonder if for solid AI we need to make ai-biology first. To be concrete about what I'm saying now, consider that we are not in conscious control of the generation of our skin cells or the processes of our kidneys, but they do organically persist on their own. Maybe we need to get ai/biology so tight-knit first, so that we can have a proper vehicle for an AI-mind to journey within, while also not necessarily rewriting the code for its heartbeat when it just wants to remember names of common phenomena.
- xapata 8y agoThat's a factor, but it doesn't explain the phenomenon that human-imperceptable transformations of an image can dramatically shift a NN's outputs.
- salawat 8y agoAgain, look to biology. That thing you are modeling. Humanly "imperceptible" is a very loaded term. Human perception has billions upon billions of networks worth of filtering going before we even boil down our environment to the "interesting" stuff. Furthermore, if you take a snapshot of that network after training, you're fit to the training data. The network has lost it's plasticity. Take a potato, put it on the ground, train the network on other potato shots. Now show it a potato shaped asteroid. Now show it a French fry. What is the potato-ness that this potato+detector is ACTUALLY homing in on? Keep in mind, this structure is trained on digital encodings of maps of light and color. The function may not be a perfect semantic detector of potato-ness. It just knows what patterns of bits MIGHT be potatoes. And when you are working on bit level encodings, one bit translates to a lot of change, even if it is imperceptible to a human looking at a rendering on a screen. Heck, there is no guarantee that the function it's emulating is well defined outside the training data set. Neural networks are GOING to be fickle. You're trying to coerce "reliable, repeatable. generifiable results" out of a simulation of the same stuff that drives five year olds and emotional people. Consider yourself lucky the program hasn't opened the CD tray and demanded you insert crayons.
- Eridrus 8y agoThis paper is definitely quite interesting, but before this thread turns into a bunch of NN hate, here is another take: "Do CIFAR-10 Classifiers Generalize to CIFAR-10?" - https://arxiv.org/abs/1806.00451 https://arxiv.org/abs/1806.00451 They use the same procedure used to construct CIFAR-10 to construct a new test set and test a bunch of state of the art results on the new test set. They see a generalization gap, but the relative order of SoTA results remains (roughly) the same. So, yes, test set validation is an overestimate of real-world performance of these systems, but progress on test sets is indicative of progress in real-world settings. And to remember this in context, no-one is asking "do traditional computer vision systems understand images", because they were explicitly just looking at image statistics.
- vowelless 8y agoHere is the new dataset: https://github.com/modestyachts/CIFAR-10.1 https://github.com/modestyachts/CIFAR-10.1
- AndrewKemendo 8y agotest set validation is an overestimate of real-world performance of these systems, but progress on test sets is indicative of progress in real-world settings I think this is important point to make and talks to the real life limitations of run of the mill SL techniques. If you can't validate with noisy "real world" data, then you're basically staying inside the box and hoping that you've curated well enough to mimic real world conditions. I still like RL system for the improvement on this, however it's much more difficult in practice.
- Cybiote 8y agoAs you point out, that the relative ordering remained stable means progress being made is not simply overfitting to the test set. The negative finding from the paper you link to is in how much accuracy drops considering how slight modifications to the test set were. In their own words: > We view this gap as the result of a small distribution shift between the original CIFAR-10 dataset and our new test set. The fact that this gap is large, affects all models, and occurs despite our efforts to replicate the CIFAR-10 creation process is concerning. > Nevertheless, the accuracy of all models drops by 4 - 15% and the relative increase in error rates is up to 3×. This indicates that current CIFAR-10 classifiers have difficulty generalizing to natural variations in image data. It remains to be seen if others can think up ways to maintain performance that the authors did not manage.
- gwern 8y agoIt's interesting that they identify striding as the culprit. Striding is also, according to some people at Google AI, the reason why VGG is one of the few good CNNs for doing style transfer (which is otherwise quite mysterious: https://www.reddit.com/r/MachineLearning/comments/7rrrk3/d_eat_your_vggtables_or_why_does_neural_style/ https://www.reddit.com/r/MachineLearning/comments/7rrrk3/d_e...).
- milani 8y agoI expected a reference to a kind of deformable convnet and its variants[1] that try to learn natural transformations. [1] https://arxiv.org/abs/1703.06211 https://arxiv.org/abs/1703.06211
- amelius 8y agoIIUC, CNNs are just a computational trick to reduce the number of parameters in the network, and train for all possible translations at once. So how would a network perform w.r.t. translations if it was expanded, i.e., topology similar to the CNN but with all parameters expanded, and trained on translated images?
- bjornsing 8y ago> Taken together our results suggest that the performance of CNNs in object recognition falls far short of the generalization capabilities of humans. No shit! An earth shattering result. :P
- John_KZ 8y agoLet me save you some time on why: >While VGG16 has 5 pooling operations in its 16 layers, Resnet50 has only one pooling operation among its 50 intermediate layers and InceptionResnetV2 has only 5 among its intermediate 134 layers. Aka being computationally cheap with pooling causes weird sampling issues. The paper has a terrible pompous title that's just wrong. Modern CNNs generalize wonderfully on small translations. They just managed to break a couple of old CNNs and go on to claim they broke AI research or something.
- microtherion 8y agoThe paper claims the opposite: "[...] jaggedness is greater for the modern, deeper, networks compared to the less modern VGG16 network. While the deeper networks have better test accuracy, they are also less invariant."
- ryanx435 8y agoBecause they are fake news * Bu dum Tish
- fallingfrog 8y agoTime shift invariance, which is used when you assume that past patterns predict future results, is one form of invariance. Space shift invariance and rotation invariance are others. You have to specifically program your nn to look for time shift invariance; stands to reason you'd want to design your architecture for the spacial invariances too. In other words: it's not reasonable to expect a neural net to derive space and rotation invariance from first principles, and I'd be very surprised to learn that the human brain didn't have special purpose hardware for accomplishing those things.
- dr_zoidberg 8y agoYou made me think of the Margaret Thatcher Illusion, for which the best example I found was Dr. Phil Plait[0]. Funny thing, the idea of Capsule Networks (Hinton et al)[1, 2] seems to tackle some of these issues, though it's they're a bit young yet and more more work. [0] https://www.opticalspy.com/opticals/dr-phil-plait-optical-illusion https://www.opticalspy.com/opticals/dr-phil-plait-optical-il... [1] https://arxiv.org/abs/1710.09829 https://arxiv.org/abs/1710.09829 [2] https://hackernoon.com/what-is-a-capsnet-or-capsule-network-2bfbe48769cc https://hackernoon.com/what-is-a-capsnet-or-capsule-network-...
- fallingfrog 8y agoI think that convolutional networks are really mostly a way to address space shift invariance too; and I even wonder if the same job could be done by some special purpose code that just runs the underlying neural net with a bunch of different rotations of the same image. That's probably how they do it now.. I feel like that's probably close to the optimal approach.
- magicalhippo 8y agoHaving dabbled with image processing, your comment reminded me of the Fourier Mellin transform, which can be used for translation, rotation and scale invariant feature detection. I did a quick search and came up with this paper[1], and while they don't use the Fourier-Mellin transform, they do use a log-polar transform[2], where rotation and scaling are transformed to translations. This, they claim, result in much improved classification of rotated and scaled data. I haven't been keeping up on the latest in the ML field tho, so maybe it's rubbish, but like you I'd be surprised if some extra processing doesn't play a role. [1]: https://arxiv.org/abs/1709.01889 https://arxiv.org/abs/1709.01889 [2]: https://sthoduka.github.io/imreg_fmt/docs/log-polar-transform/ https://sthoduka.github.io/imreg_fmt/docs/log-polar-transfor...
- klausjensen 8y agoFor those who (like me) did not know what CNN is: In machine learning, a convolutional neural network (CNN, or ConvNet) is a class of deep, feed-forward artificial neural networks, most commonly applied to analyzing visual imagery. (Source: https://en.wikipedia.org/wiki/Convolutional_neural_network https://en.wikipedia.org/wiki/Convolutional_neural_network)
- d--b 8y agoPerhaps, I'm saying just perhaps, there is a reason why humans are good at discerning things in the very narrow field of vision that's straight ahead, and not very good at the peripheral vision. Pointing first and then classify may help in solving those issues. Maybe?
- PeterisP 8y agoThe reason why humans are good at discerning things in the very narrow field of vision that's straight ahead, and not very good at the peripheral vision is biological, there's simply a much lower density of receptors ("pixels") in the periphery, so in the periphery there's much less sensory information for the brain to work with.
- zer0faith 8y agoI thought this reference was for CNN News Network. Silly me..
- candiodari 8y ago"How do humans do small image transformations ?" Perhaps the answer is simple : REM (the awake variant), and https://www.youtube.com/watch?v=quJEyTvDdfY https://www.youtube.com/watch?v=quJEyTvDdfY
- candiodari 8y agoIt's funny but looking at those prediction graphs bouncing up and down like crazy ... one immediately thinks "yep I know, that's exactly what they said would happen if I used polynomials for fitting functions". And yes, that's what polynomials do. They don't have good local behavior: if a -> f(a) and b -> f(b) are fitted with a higher-degree polynomial then the image of f between a and b will often be infinite (meaning there is some value x between a and b where f(x) is infinite), starting at very high degrees (ie. deep networks) there will probably be many such points. Intuitively I think of it like this: if you look at the "real world" as a function you can make a couple of very general observations. F(x) -> doesn't have very much information about the world, and it's very hard to make sense of. Most of it just doesn't seem relevant. d/dx F(x) ... much more relevant. d^2/dx F(x) ... also pretty interesting. d^3/dx F(x) less interesting but occasionally important. d^4/dx F(x) ... nobody cares (also if you take camera images and calculate this, it'll be almost exclusively zeroes). Secondly there are strong "domains" in the real world that we just seem to be unwilling to accept. Polynomials are good in the sense that if you get the equation for a stone dropping onto your foot really, really tighly correctly fitted, that equation holds up for the movement of an entire planet, which is great. But why bother ? It is much more valuable to be able to predict whether a stone will fall on my foot than how Venus will move. That's if you get it right. If you get the polynomial degree of your equation wrong ... it makes utterly ridiculous predictions. Many other approximation methods don't suffer from this problem Doesn't happen with spline, beziers, even taylor approximations have better behavior.