3 ms·
Lecun has identified a real problem for AI -- the need to understand the real world, the link between intelligence and prediction over time. But the tools he i
by pakl 10y ago
Lecun has identified a real problem for AI -- the need to understand the real world, the link between intelligence and prediction over time. But the tools he is using are not the right ones.
Deep conv nets were not designed with prediction over time in mind. Here's one reason: Deep convolutional nets do not handle dynamical information from their lowest layers. By design, conv layers and pooling layers immediately begin discarding spatial information that could be used for building up predictions.
In contrast it is possible to start with recurrent/feedback networks at the very first layers of the network. These initial layers can begin building up predictions at the pixel, color, and lighting level (example: our recent preprint[1]).
My colleague, who as some of you know enjoys blogging, wrote a more thorough post in response to LeCun's recent CMU lecture on the same topic as these slides[2].
[1] https://arxiv.org/abs/1607.06854 https://arxiv.org/abs/1607.06854
[2] Blog: "A few comments on the Yann LeCun lecture at CMU, 11.2016" http://blog.piekniewski.info/2016/11/21/yann-lecun-cmu-11-2016-comments/ http://blog.piekniewski.info/2016/11/21/yann-lecun-cmu-11-20...
- felippee 10y agoHi, I'm the blogger. I just want to add a very simple statement: In order to create a model of the world, the machine learning substrate has to have the capacity to cover the observed dynamics. Hard not to agree with this, almost sounds like a tautology. Now let's try to derive conclusions: - World dynamics is full of multi-scale interactions (e.g. in vision illumination of a single pixel depends on the whole scene and the whole scene depends on many tiny details). To capture that, the machine learning substrate has to allow low level representations access high level stuff. Hence feedback all over the place. Which is exactly what is seen in the biological cortex. This is not recurrent layer made of LSTMs, this is a FULLY RECURRENT system. This will not be achieved with any franken-neocognitron deep network neither with MSE, nor adversarial nor even triple adversarial loss function. This requires a new approach and together with several colleagues after a few years of continuous and intense thinking and modelling we have proposed a solution: http://blog.piekniewski.info/2016/11/04/predictive-vision-in-a-nutshell/ http://blog.piekniewski.info/2016/11/04/predictive-vision-in... As well as a full paper https://arxiv.org/abs/1607.06854 https://arxiv.org/abs/1607.06854 Now, I'm not saying this is all done. This is just a beginning of a really exciting research adventure and many things look very promising. It will require however for the AI field to get out of a pretty "deep" local minimum it is in right now.
- eli_gottlieb 10y agoYeah, not to be too dismissive or cliched, but these slides come off as if LeCun just now got around to reading some Andy Clark or reading neurosci/cogsci papers by Tenenbaum or Friston. The idea that the human brain has to work via active prediction rather than passive signal processing has been well-established in cogsci and neurosci for a while now. The interesting question, then, is how we make machine-learning systems do prediction well. Probabilistic/generative models have been "wandering in the desert" for a while now because Monte Carlo methods are just so slow, especially for high-dimensional, hierarchical prediction problems like we want to solve in machine learning. On the upside, STAN now has automatic variational inference for continuous probability models, and work on things like the "concrete distribution" (https://arxiv.org/abs/1611.00712 https://arxiv.org/abs/1611.00712) can help us continuously approximate discrete probabilistic reasoning. Maybe as these techniques move into the mainstream in systems like Venture or Picture we can start to scale up predictive/generative/probabilistic modelling to match optimization-based connectionist methods?
- felippee 10y agoThe idea is not new indeed, Andy Clarks's review paper was very inspiring to us. But it has not been detailed to a point of implementation/scaling. PVM is an attempt to implement it in a "connectionist" way, but frankly all I need are associative memories, and how they are implemented I don't care. So a "probabilistic PVM" is totally feasible. In fact we discuss in the paper various possibilities in which a PVM like meta-architecture can be implemented.
- ceejay 10y agoWould it be feasible to replace the "common" component of a recurrent neural network with a convolutional neural network? My lay person's impression is that at its most basic level a recurrent neural network is simply a "conveyor belt" of neural nets which are affected by external weights as well as by the weights from within the network. More precisely the "internal" weights coming from the layer of perceptrons operating at 1 level shallower than itself. So we're dealing in essence with 2 dimensions (shallower to deeper, and older to newer) instead of just one (shallower to deeper).
- felippee 10y agoWhy would you want that? Conv net is not some magic. It's just a crude way to reduce dimensionality by loosing spatial location. For some things it works, for some it doesn't. I think the shift in thinking should rather be: instead of trying to build the best possible associative memory to associate some A with some B, take the memory modules we have (perhaps not perfect) and try to build something bigger out of them. A dynamical model of the observed reality seems like a great thing to build out of such modules. And this is what the PVM is. Currently made out of shallow, plain vanilla perceptrons, builds a structure which can be arbitrarily deep. Without any "magical" tricks such as dropout, relu, convolution, pooling etc.