8 ms·
Failures of Deep Learning
- smdz 10y agoLink to the PDF: https://arxiv.org/pdf/1703.07950.pdf https://arxiv.org/pdf/1703.07950.pdf
- dmreedy 10y agoI think some of the most exciting and interesting work comes out of proving, not just capabilities, but constraints for systems, be it Gödel, Shannon, Aaronson, or any of the others in the smaller-than-desirable tradition of those who say, "No". I think a better understanding what Deep Learning can't do (well) is fertile material for better understanding the kinds of problems it can do, and am very excited to see more work in this space, and movement towards an underlying structural theory.
- bitL 10y agoI think the issue with Deep Learning is that it is like a hyperdimensional optimization heuristics sequence surpassing what human mind can comprehend and pushing limits of computing (depth of 100 layers max at the moment). Given how difficult are far more trivial optimization techniques and proving bounds for them, it seems the times where we could just define a new approach and prove some nice properties are behind us :-(
- boxcardavin 10y agoI don't think the problem here is what the human mind can comprehend because identifying where these methods break down is actually pretty easy from a math perspective, and the breakdown doesn't change as you go up and down with the number of dimensions. What has always surprised me about ML and DL is how far we can stretch simple regression techniques and how useful it is on such a variety of problems. I agree that it has turned into an NP problem in a lot of cases now but the concepts behind the problem of N-dimensional optimization are pretty well understood.
- rectangletangle 10y agoA lot of useful analogies for N-dimensional space can be abstracted from the simple transition from 2 to 3 dimensions.
- joeyo 10y ago"To deal with a 14-dimensional space, visualize a 3-D space and say 'fourteen' to yourself very loudly." -- Geoff Hinton
- bitL 10y agoA math student and an engineering student went together to a presentation called "14-dimensional space topology". As the presentation progressed, engineering student became more and more frustrated as she could only grasp a few simple initial examples but the deeper the presentation went, the worse was the understanding. Yet the math student was absolutely enthusiastic, often asking the presenter various complicated questions and looking he enjoyed himself a lot. With a massive headache in the end, the engineering student turned to math student and asked: "How could you understand the presentation? I was able to understand 2D, 3D but once we increased number of dimensions I got lost and never understood anything about 14D!" The math student looked at her with a confident careless look and said: "It was simple. I imagined everything in N dimensions, and then just reduced it to 14".
- wnoise 10y ago"and then let N=14" is how I've heard it.
- curuinor 10y agoIt's not like we don't know anything about NP complete problems, neither. Critical phase transition in transition of alpha on random kSAT and other stuff, the realization that this is neither necessary or sufficient for NP completeness just a really common complementary phenomenon, etc etc
- LeanderK 10y agoI strongly dissagree. Broken down, Neural Networks are surprisingly simpel and elegant. The combination makes them so powerful. Ideally, an elegant theory would abstract over these to provide bounds for arbitrary combinations, or groups of combinations.
- curuinor 10y agoWhy can't we pull out the statistical mechanics? You can't understand the 10^27 variables in a lump of coal, either, but you can still do thermodynamics. People already do it, Ganguli, some stuff from Bengio
- bitL 10y agoI wish I were so optimistic - if you use statistics, then you might remove most of interesting outliers and study just those more stable prevailing characteristics and not the problematic ones, like when using Hooke's law for designing buildings while ignoring its limited scope. And physics hits hard math limits while studying dynamic systems (i.e. fractals) anyway. We can always use Ramsey Theory stating there is no chaos possible and you can always find some regularity in anything, yet if you need to scan 10^132 items to find the order it becomes quickly uninteresting.
- curuinor 10y agoRNN is natively and obviously dynamical but FFNN can also be seen as transient dynamical system. Little Ramsey theory ever used in neural network land, but theory of dynamical systems used from quite a long time ago, from Werbos-BPTT which was explicitly envisioned as a dynamical algorithm to the description of the dynamical instability of gradients in RNN. The statistics of statistical mechanics don't necessarily need to be CLTish sorts of things, you know?
- bitL 10y agoPossibly ;-) I'll study these things in detail soon (hopefully), so far just practical experience with all funny things from DL like self-driving cars, composing music with RNNs etc. but having extensive operations research exp in the past I tend to be careful.
- tianlins 10y agoThe high dimensionality is merely for the convenience of optimization, so that we get smoother manifold on which a good approximation solution can be found. In that regard, it is interesting to think about of space of models that we humans can interpret, and the possibility of "distill" deep neural nets into models in that space.
- xamuel 10y agoThe paradox of the heap[1] shows that vagueness is an inherent, inescapable part of reality. Machine learning would provide a way around that paradox: a heap is what a neural network says is a heap! If this worked, it would be too good to be true. Thereore, it cannot work. [1] https://en.wikipedia.org/wiki/Sorites_paradox https://en.wikipedia.org/wiki/Sorites_paradox
- westoncb 10y ago> The paradox of the heap[1] shows that vagueness is an inherent, inescapable part of reality. Or it's an inescapable part of how we interface with reality, i.e. it's an artifact of the structure of the human brain.
- curuinor 10y agoBak of the Bak-Tang-Wiesenfeld model called his model the sandpile, basically explicitly to remind one of the Sorites paradox. Interesting thing to talk about. You can definitely take it as a sort of different descendent of the Ising model: it has had much less influence on the world compared to backpropagation neural net. The dynamical nature of the definition stays, which has bedraggled Sorites people for millenia, but McClelland has been talking about the dynamical nature of representation for decades too. I think some of the anti-neural net cognitive science people (Pinker, Fodor maybe? I forget) brought up some Sorites arguments when fighting the 80's connectionists.
- curuinor 10y agoWell, they already had a big instance of that, in the Minsky Papert book "Perceptrons", talking about linear separability. Talking about backpropagation (as opposed to delta rule) in opposition to that is a sort of mushing of history but it's interesting to think about
- deleted 10y ago[deleted]
- bra-ket 10y agothe biggest failure of deep learning is the lack of common sense
- visarga 10y agoIt is growing, gradually: ontologies, word embeddings, mechanical dynamics prediction, but it will take some time. I don't know either why there isn't more of a push to bring together all the common sense resources.
- Houshalter 10y agoJust yesterday there was a big discussion on here about how academic papers needlessly complicate simple ideas. Mainly by replacing nice explanations with impenetrable math notation in an attempt to seem more formal. This paper is very guilty of this. E.g. page 5. They attempt to explain a really simple idea, that they generated images of random lines at a random angle. Then labelled the lines positive or negative examples, based on whether the angle was greater than 90 degrees or not. Then they take sets of these examples. And label them based on whether they contain an even or odd number of positive examples. They take several paragraphs over half a page to explain this. Filled with dense mathematical notation. If you don't know what symbols like U, :, ->, or ~ mean, you are screwed because that's not googleable. It takes way longer to parse than it should. Especially since I just wanted to quickly skim the ideas, not painfully reverse engineer them. Hell, even the concept of even or odd, is pointlessly redefined and complicated as multiplying + or - 1's together. I was scratching my head for a few minutes just trying to figure out what the purpose of that was. It's like reading bad code without any comments. Even if you are very familiar with the language and know what the code does, it takes a lot of effort to figure out why it's that way. If it's not explained properly. The worst part is, no one ever complains about this stuff because they are afraid of looking stupid. I sure fear that by posting this very comment. I actually am familiar with the notation used in this example. I still find it unnecessary and exhausting to decode.
- abecedarius 10y agoI've complained about this sort of thing too (usually more privately). The upside is supposed to be precision and concision, but as you point out in this case they spent all their concision on buying precision. For this precision I think code should be considered more often instead of the usual informal math notation, which sometimes gets sloppy and harder to follow from outside the author's research community. (Edit: I'm not accusing that passage of sloppiness, but it's a problem that's frustrated me before with math in place of code.)
- nojvek 10y agoI will dance around the day someone figures our an AI to translate papers into readable code that compiles into a nice jupyter like notebook format. Some papers are quite easy to comprehend and truly ground breaking e.g the CNN paper by Alex and Hinton. Some like this are like "WTF are you even trying to say?"
- csfoo 10y agoThe lead author will be giving a talk on this work next week (which will be live streamed and recorded) as part of a workshop on Representation Learning: https://simons.berkeley.edu/talks/shai-shalev-shwartz-2017-3-28 https://simons.berkeley.edu/talks/shai-shalev-shwartz-2017-3...