4 ms·
How is "deep learning" different from "neural network"?
by stewie2 14y ago
How is "deep learning" different from "neural network"?
- dave_sullivan 14y agoWell--two answers: 1) It's not. It's just a buzz word that people are going to use to separate the current (2006+) research from older research concerning neural networks. This is to draw a clear (and potentially self serving?) distinction between old neural networks that were discredited due to their lack of results vs current research that produces much much better results. So, it's neural networks rebranded. Oh my. 2) "traditional" neural networks and what people are using now are very different--mostly because what people are doing now actually works. Deep learning refers to deep neural networks, which take more traditional neural networks and stack them on top of each other to form a hierarchy of representations that ends up being effective for all kinds of stuff. Not that deep neural networks are a new concept--the newness is more that this is now practical rather than theoretical. So, deep learning is like 20% bullshit, 80% the real deal. Still lots of work to be done, but I think "deep learning" is a nice buzzword to describe the current state of the art as far as neural networks go. It's all neural networks--but this time it's different, haha.
- Yoshua 14y agoThe idea of having multiple levels of representation (deep learning) goes beyond neural networks. A good example is the recent work (award-winning at NIPS 2012) on sum-product networks, which are graphical models whose partition function is tractable by construction. Several important things have been added since 2006 (when deep learning was deemed to begin) to the previous wave of neural networks research, in particular powerful unsupervised learning algorithms (which allow very successful semi-supervised and transfer learning - 2 competitions won in 2011), often incorporating advanced probabilistic models with latent variables, a better understanding (although much more remains to be done) of the optimization difficulty of training gradient-based systems through many composed non-linearities, and other improvements to regularize better (such as the recent dropouts) and to rationally and efficiently select hyper-parameters (random sampling and Bayesian optimization). It is also true that sheer improvements in computing power and amounts of training data are in part responsible for the impressively good results recently obtained in speech recognition (see recent New York Times article, 24 nov, J. Markoff) and object recognition (see NIPS 2012 paper by Krizhesky et al).