6 ms·
Deep Learning 101
- Noxchi 13y agoWhat does it take to be good at machine learning such as this? In terms of mathematics, and computer science knowledge? I know how to code through self learning, and I've pretty much solely done web development. So I barely know much CS. Also not very good at math. So what are the essential prerequisites you would say are necessary for doing neat, useful stuff with machine learning?
- mbeissinger 13y agoDefinitely a solid foundation in linear algebra and statistics (mostly Bayesian) are necessary for understanding how the algorithms work. Check out the wiki portals (http://en.wikipedia.org/wiki/Machine_learning http://en.wikipedia.org/wiki/Machine_learning) and (http://en.wikipedia.org/wiki/Artificial_intelligence http://en.wikipedia.org/wiki/Artificial_intelligence) for overviews of the most common approaches. Also, Andrew Ng's coursera course on machine learning is amazing (https://www.coursera.org/course/ml https://www.coursera.org/course/ml) as well as Norvig and Thrun's Udacity course on AI (https://www.udacity.com/course/cs271 https://www.udacity.com/course/cs271)
- elq 13y agoMath. math. math. Linear algebra. Bayesian statistics. MUST know these inside out, upside down. Vector calculus. Convex optimization. A boatload of machine learning literature. The ideas coalescing into deep learning are based more than a decade of research. If you know nothing about math... I can't imagine getting to the point of understanding deep learning (which is a fairly rapidly evolving area) without at least 2-3 years of very hard work. This class is a reasonable attempt to give a quick intro to one major source for DBNs https://www.coursera.org/course/neuralnets https://www.coursera.org/course/neuralnets understanding this course is a good benchmark.
- sown 13y agoI ran through Udacity's CS373 course first (https://www.udacity.com/course/cs373 https://www.udacity.com/course/cs373). It was neat.
- arjunrajjain 13y agoThis is awesome!
- mbeissinger 13y agoThanks!
- cocoflunchy 13y agoOT but the text selection behavior on this page is fascinating! (Or horrific if you don't want to be nice). I've never seen anything like it. https://www.dropbox.com/s/4k72g8b2tl3mgzt/Screenshot%202013-11-13%2020.29.48.png https://www.dropbox.com/s/4k72g8b2tl3mgzt/Screenshot%202013-...
- mbeissinger 13y agoAhh what are you viewing it on?
- cocoflunchy 13y agoChrome 31.0.1650.48 on OSX
- sudont 13y ago31.0.1650.48 here as well. It appears to be a bug with a combination of the ::selected pseudoelement in conjunction with the font Georgia. My guess is it's a Chrome-on-mac bug (Firefox is fine), not a site coding error. Disabling either the font or the selection style fixes it. Most likely a text rendering issue. At work we've noticed Chrome getting buggier in relation to that, as well as retaining DOM node properties via redraws.
- mbeissinger 13y agoOdd. I can't replicate on any browser but I don't have OSX. Does it work fine on your other browsers?
- ingrownpsyche 13y agoThat's fine for me, but I'm getting ligatures for every st. I thought it was deliberate (and a little pretentious really) but you're not getting them so hurray webfont or something I suppose.
- brandonb 13y agoThis is a cool tutorial! It's ironic that deep neural networks have become the biggest machine learning breakthrough of 2013: they were also the biggest machine learning breakthrough of 1957. The idea dates back to the Perceptron, one of the oldest ideas in AI. One thing to note: although there was a lot of initial excitement about Restricted Boltzman Machines, Auto-encoders, and other unsupervised approaches, the best results in the last year or so have all used conventional the back-propagation algorithm from 1974, with a few tweaks. http://en.wikipedia.org/wiki/Backpropagation http://en.wikipedia.org/wiki/Backpropagation Ben Lorica wrote a good article on the latest deep learning research from Google, and what's changed since neural networks were last popular in the 1980's: http://strata.oreilly.com/2013/10/deep-learning-oral-traditions.html http://strata.oreilly.com/2013/10/deep-learning-oral-traditi... What's old is new again.
- seiji 13y agoRBMs and auto-encoders use backprop too. They just use it for "fine tuning" (due to running pre-training first) instead of propagating error derivatives from randomly initialized weights. (Thus concludes the smartest thing I've said all day.)
- mbeissinger 13y agoYep the ideas from the 50's have definitely reappeared now we have the compute power and methods to implement them at a large scale. That article gives a nice perspective. One of the best breakthroughs has been this notion of layer-wise pretraining, which allows the backpropagation algorithm to not get stuck in local minima so easily. It provides a good guess to the starting starting points for the weights. Otherwise, the biggest issue with backpropagation historically has been the diffusion of weights as the layers increase; it is hard to attribute the causality or what portion of the update weighting should be applied to each node since it grows exponentially. This pretraining idea helps against that.
- brandonb 13y agoThat's what I thought too! But according to my friends on the Google Brain team, unsupervised pretraining is now thought to be an irrelevant detour. In 2006, Hinton introduced greedy layer-wise pretraining, which was intended to solve the problem of backpropagation getting stuck in poor local optima. The theory was that you'd pretrain to find a good initial set of connection weights, then apply backprop to "fine-tune" discriminatively. And the theory seemed correct since the experimental results were good: http://www.cs.toronto.edu/~hinton/absps/fastnc.pdf http://www.cs.toronto.edu/~hinton/absps/fastnc.pdf http://machinelearning.wustl.edu/mlpapers/paper_files/NIPS2006_739.pdf http://machinelearning.wustl.edu/mlpapers/paper_files/NIPS20... Does pretraining truly help solve the problem of poor local optima? In 2010, some empirical studies suggested the answer was yes: http://machinelearning.wustl.edu/mlpapers/paper_files/AISTATS2010_ErhanCBV10.pdf http://machinelearning.wustl.edu/mlpapers/paper_files/AISTAT... But that same year, a student in Geoff Hinton's lab discovered that if you added information about the 2nd-derivatives of the loss function to backpropagation ("Hessian-free optimization"), you could skip pretraining and get the same or better results: http://machinelearning.wustl.edu/mlpapers/paper_files/icml2010_Martens10.pdf http://machinelearning.wustl.edu/mlpapers/paper_files/icml20... And around ~2012, a bunch of researchers have reported you don't even need 2nd-derivative information. You just have to initialize the neural net properly. Apparently, all the most recent results in speech recognition just use standard backpropagation with no unsupervised pretraining. (Although people are still trying more complex variants of unsupervised pretraining algorithms, often involving multiple types of layers in the neural network.) So now, after seven years of work, we're back where we started: the plain ol' backpropgation algorithm from 1974 worked all along. This whole topic is really interesting to me from a history of science perspective. What other old, discarded ideas from the past might be ripe, now that we have millions of times more data and computation?
- shon 13y agoGoogle, Twitter, Netflix, Yelp, Pandora and more are speaking on Deep Learning and RecSys this Friday at MLconf in San Francisco. We're trying to get a streaming solution going as well for those who can't make it. http://mlconf.com http://mlconf.com DISCLAIMER: This is my event
- mbeissinger 13y agoI'll definitely check this out if you get a stream going.
- shon 13y agoCheck mlconf.com on Friday. The default will be Ustream here: http://www.ustream.tv/search?q=mlconf http://www.ustream.tv/search?q=mlconf We'll also post that and any updates to the main site on Friday.
- luu 13y agoPersonally, I've found that I don't retain much of this sort of material without working through exercises. If you learn the same way, you might want to check out the series of progressive exercises from Andrew Ng here: http://ufldl.stanford.edu/wiki/index.php/UFLDL_Tutorial http://ufldl.stanford.edu/wiki/index.php/UFLDL_Tutorial For reference, I have a copy of my solutions here: https://github.com/danluu/UFLDL-tutorial https://github.com/danluu/UFLDL-tutorial. Debugging broken learning algorithms can be tedious in a way that's not particularly educational, so I tried to find a reference I could compare against when I was doing the exercises, and every copy I found had bugs. Hope having this reference helps someone.
- mbeissinger 13y agoNice!
- misiti3780 13y agoawesome - thanks for the links
- msvan 13y agoFor some more elementary material, I also recommend Andrew Ng's machine learning course on Coursera. He's a great teacher.
- kot-behemoth 13y agoLink for the impatient https://www.coursera.org/course/ml https://www.coursera.org/course/ml Looks great indeed!
- msvan 13y agoFor some more elementary material, I also recommend Andrew Ng's machine learning course on Coursera. He's a great teacher.
- dave_sullivan 13y agoThis is a really good write up. For people looking for practical experience with these types of methods, I'd also recommend checking out theano and/or pylearn2 (which is built w/ theano). theano: http://deeplearning.net/tutorial/ http://deeplearning.net/tutorial/ pylearn2: http://deeplearning.net/software/pylearn2/ http://deeplearning.net/software/pylearn2/ NIPS, a big ML conference, is in December, so expect to see a large amount of new ideas and applications re: deep learning to come out of that.
- mbeissinger 13y agoThanks for putting those links up - the Theano documentation has some great tutorials for how to code these in practice.
- sidcool 13y agoAn excellent 101 article.
- mbeissinger 13y agoThanks!
- ma2rten 13y agoI've followed the developments in Neural Networks somewhat, but have never applied deep learning so far. This is seems like a good place to ask a couple of question I've been having for a while. 1. When does it make sense to apply deep learning? Could it potentially be applied successful applied to any difficult problem given enough data? Could it also be good at the type of problems that Random Forest, Gradient Boosting Machines are traditionally good at versus the problems that SVMs are traditionally good at (Computer Vision, NLP)? [1] 2. How much data is enough? 3. What degree of tuning is required to make it work? Are we at the point yet where deep learning works more or less out the box? 4. Is it fair to say that dropout and maxout always work better in practice? [2] 5. What is the computational effort? How long e.g. does it take to classify an ImageNet image (on a CPU / GPU)? How long does it take train a model like that? 6. How on earth does this fit into memory? Say in ImageNet your have (256 pixels * 256 pixels) * (10,000 classes) * 4 bytes = 2.4 GB, for a NN without any hidden layers. [1] I am overgeneralizing somewhat, I know. It's my way to avoid overfitting. [2] My lunch today was free.
- dwiel 13y agoI don't have great answers to the other questions, though I too am interested in them. #5) [1] has a some python code and timings mixed in to the docs. One such example (stacked denoising autoencoders on MNIST): By default the code runs 15 pre-training epochs for each layer, with a batch size of 1. The corruption level forthe first layer is 0.1, for the second 0.2 and 0.3 for the third. The pretraining learning rate is was 0.001 and the finetuning learning rate is 0.1. Pre-training takes 585.01 minutes, with an average of 13 minutes per epoch. Fine-tuning is completed after 36 epochs in 444.2 minutes, with an average of 12.34 minutes per epoch. The final validation score is 1.39% with a testing score of 1.3%. These results were obtained on a machine with an Intel Xeon E5430 @ 2.66GHz CPU, with a single-threaded GotoBLAS. #6) The size of the NN is not typically num_features * num_classes, but rather num_features * num_layers where num_layers is commonly 3-10 or so. If you want a (multi-class) classifier, you first feed your neural network a bunch of examples, unsupervised. Then once you've got your NN built, you feed the outputs of the NN to a classifier like SVM or SGD. The idea is that the net provides more meaningful features than you would have if you used hand crafted features or the raw input data itself. [1] http://deeplearning.net/tutorial/SdA.html#sda http://deeplearning.net/tutorial/SdA.html#sda
- faxmulder 13y agoVery interesting stuff written in a clear way. I'm actually finishing my master thesis on music genre recognition through machine learning, which is focused more on traditional ensemble learning, but I think that it would be nice to study deep learning in greater detail. Thanks!
- mbeissinger 13y agoAwesome, do you have any demos for the music genre recognition?
- faxmulder 13y agonot yet, I've still some work to do. One question: do you think that Optimum-Path Forests could be used also in the context of deep learning?
- albertzeyer 13y agoCan anyone recommend a good book on the topic? And maybe other recent neural network research topics; I'm esp. also interested in recurrent networks like LSTM.