6 ms·
What's Wrong with Deep Learning?
- codewithcheese 11y agoWow what in incredible amount of knowledge in those slides. Is there a video of the keynote?
- sjtrny 11y agoCVPR usually makes all talks available when the conference has concluded. Watch this page http://www.pamitc.org/cvpr15/ http://www.pamitc.org/cvpr15/.
- jhartmann 11y agoIt is always impressive to me how both Professor LeCun, Bengio & Hinton stuck to their guns and worked on these problems while others were not so interested. I'm very glad Hinton's group at DNN Research was able to blow away the competition in the Imagenet challenge. Now many more people are working on these ideas and some really amazing things have been accomplished in a very short time. I can't wait to see where we are in five years, and I love that LeCun's points out the areas where we should focus and what we are not good at yet.
- fizixer 11y agoDo you have any information about what were the hot topics just before this new resurgence of NNs (meaning around late 90's, and early 2000's). I know symbolic AI was big in 60s, and 80s, but not sure about recent past.
- albertzeyer 11y agoSupport Vector Machines was one of the hot topics.
- oergiR 11y agoProbabilistic models. Recent research often focuses on Bayesian models. Probabilistic models have never really gone away. This presentation by LeCun actually suggests embedding neural networks inside of various types of probabilistic models: factor graphs and conditional random fields. This is, for example, how speech recognition works: the output of a neural network is fed into a probabilistic model (a hidden Markov model).
- jhartmann 11y agoActually, state of the art speech recognition has switched over to having a Recursive Neural Network directly run over the audio input. Take a look at the paper at http://arxiv.org/abs/1412.5567 http://arxiv.org/abs/1412.5567 and http://usa.baidu.com/deep-speech-lessons-from-deep-learning/ http://usa.baidu.com/deep-speech-lessons-from-deep-learning/ However combining learning features with other systems is a very powerful approach and combining SVM's on top of the learned features of a Neural Network I would say is common. I personally am more interested in approaches like Deep Fried Convnets (http://arxiv.org/abs/1412.7149 http://arxiv.org/abs/1412.7149) that combine kernel methods as part of the Neural Networks themselves.
- oergiR 11y agoI know that Andrew Ng and colleagues say that they don't use HMMs. I haven't spoken with them (I haven't seen them at speech conferences) so I do not know whether they actually believe this themselves. I believe the best comparison between "CTC" (which is billed as recurrent neural networks without the HMMs) and the traditional approach is by people at Google, Sak et al, "Learning Acoustic frame labeling for speech recognition with recurrent neural networks", ICASSP 2015. (I can't find a PDF online.)
- agibsonccc 11y agoNot to nitpick. I just want people to realize there are actually recursive nets that rely on a parser to be built (this is the recursive net that relies on backpropagation through structure). Then there is the recurrent net (LSTMs,multimodal) that rely on backpropagation through time. Talking to some of the users of Recursive nets, they will be renaming them to tree rnns which should help clear up confusion a bit.
- bra-ket 11y agokernels
- discardorama 11y agoAs far as I remember, everyone was jumping on top of SVMs and Markov Models.
- pigscantfly 11y agoIn the case of vision, there were a lot of things going on, but support vector machines, sliding window search, descriptors with invariances to different transforms like SIFT and HOG, spatial pyramids, deformable parts models, and mixtures of gaussians are all hot topics that you'll regularly see in papers from the early 2000's. There was a lot of work on improving runtime for these techniques going on, since training SVM experts and evaluating anything over a sliding window search space is very expensive. You'll also see a lot of work on approaches rooted in graph theory like conditional random fields, different flavors of Markov models, min-cut/max-flow based algorithms, and other types of probabilistic graphical models. Most of this stuff is still in widespread use; AI is a big field.
- deleted 11y ago[deleted]
- spin 11y agoI really like the book "Pattern Recognition and Machine Learning" by Christopher Bishop. It's packed full of all the latest-and-greatest algorithms. WRT your question, an interesting "feature" of that book is that it was published just before deep neural networks started taking off, so there's no mention of DNNs in the book. You can see what the world was like right when they started taking off.
- chestervonwinch 11y agoI would like to see or hear more regarding the theory slides - in particular on the objective being a piecewise polynomial, and the distribution of weights using random matrix theory. Anyone know where I could find more?
- selimthegrim 11y agohttp://arxiv.org/abs/1412.0233 http://arxiv.org/abs/1412.0233 (and the references within he cites by Gerard Ben Arous)
- deleted 11y ago[deleted]
- aswanson 11y agoIs there a volume summary of research papers or recommended book(s) for DNN?
- Animats 11y agoIs that document available in some standard format? The player that's playing it from Google Docs is buggy, and about 20% of the slides display an error message. "10:01:47.662 Cross-Origin Request Blocked: The Same Origin Policy disallows reading the remote resource at https://drive.google.com/viewerng/img?id=ACFrOgBySwSrGvI-XLLt8yzMipqv7ffA6jN_fFpTRl88JGL9KzU39f4S0oroqq45kVAZ7staj8oYE6yPEsqdHD6S8r09_m-ZRpYjNTUwNlILDrA2-H463PBLY9LZSb8=&w=2000&page=46 https://drive.google.com/viewerng/img?id=ACFrOgBySwSrGvI-XLL.... (Reason: CORS header 'Access-Control-Allow-Origin' missing).1 <unknown> "
- ipsin 11y agoLook to the nav bar at the top of the page. The source file is (purportedly) a PDF and there's a download button.
- discardorama 11y agoIf someone were to ask me what's wrong with DL (and not that anyone would, since I'm an unknown), I'd say the lack of theory. Most DL results look very hacky to me. Someone says Max Pooling works; someone again comes along and says it's not necessary. Someone says sigmoid or tanh are the best activation functions; someone else says ReLUs are better. And so on. Why? Why is one better than the other? I'm no biologist, but I don't think our brains are going around trying to do a grid search for the best hyperparameters. Most DL results today are the result of throwing 1000s of Titans on the problem and then sitting back for a week for the beast to cough up a solution. Tangential nitpick: one (very minor) nit I have with Prof LeCun's presentations is that I don't see him give more credit to Hinton and Schmidhuber. Hinton is mentioned a couple of times (3), but Schmidhuber is totally ignored; for example, when he mentions LSTM, it's cited as [Hochreiter 1997], even though it was a join publication with Schmidhuber. It should be cited as [Hochreiter et 1997], as he does in the very next line.
- ma2rten 11y agoLack of theory is actually mentioned as one of the issues in the presentations. I don't think your examples are good though, Max polling reduces noise. RuLU learn faster than Sigmoid or tanh.
- kmicklas 11y ago> I don't think your examples are good though, Max polling reduces noise. RuLU learn faster than Sigmoid or tanh. That's not theory, that's just observation of the results. Why should we expect it to work that way?
- dave_sullivan 11y agoI agree re: lack of theory. No easy answer on that one other than "keep looking and get more people to help". We are making major practical gains along the way (although many are quick to discount those--"that's it???"). It's science in practice, theory follows. I disagree re: the 1000s of titans thing. Google, Baidu, etc are building large GPU clusters and have basically shown "similar resources = similar results", but everyone else is mostly using single machine--maybe multi-GPU--and doing fine. A single Titan X is a BEAST for deep learning--nobody is using 1000s and you only need 1 for great results on most datasets I've seen. On the subject of Schmidhuber, I saw him speak once and he spent half the talk explaining how he invented everything he's talking about (EVERYTHING!) and the other half talking about how no one gives him credit. I'm half joking, but I think there's more to his story. Or it's a miscarriage of justice.
- Animats 11y agoSee slides 134-135. It's amazing that works. They get induction without any understanding at all. It still gets the right answer. Intelligence may be dumber than we thought it was.
- sgt101 11y agoThe final 20 or so slides about building general AI using deep learning strike me as really interesting. Seymour Papet said that if you can fit concepts into your cognitive architecture then they are learned. I think that this part of the presentation speaks to a need to demonstrate this as "learning" proper. It's strange, because I believe that this needs to happen, but the idea that you would do it with an "all nn all the way down" architecture, rather than breaking out into a symbolic layer a-la SOAR just seems odd.
- kragen 11y agoFor some reason, most of these slides say, "¡Vaya! Hubo un problema para cargar la página." Is there a better URL, maybe with the PDF itself? (https://doc-04-4c-docs.googleusercontent.com/docs/securesc/ha0ro937gcuc7l7deffksulhg5h7mbp1/0ptf7tt7v8qkg9j40pi37h2cvhm14jn8/1434326400000/06392201561352539427/*/0BxKBnD5y2M8NVHRiVXBnOVpiYUk?e=download https://doc-04-4c-docs.googleusercontent.com/docs/securesc/h... doesn't look like it's going to work reliably for other people.)
- abecedarius 11y agoI downloaded the pdf; will mail you.
- rlucente 11y agoI have attempted to put together the math stack for deep learning at http://rlucente.blogspot.com/2014/08/deep-learning-mathematical-stack.html http://rlucente.blogspot.com/2014/08/deep-learning-mathemati...