4 ms·
Since you seem to know this stuff, I hope you doing mind if I ask you a marginally related question... It seems like deep learning is based on greedily buildin
by lliiffee 16y ago
Since you seem to know this stuff, I hope you doing mind if I ask you a marginally related question... It seems like deep learning is based on greedily building representations in an unsupervised setting. (Not sure if this is a defining characteristic of deep learning or just a common usage.) The postulate seems to be that greedily learning in a supervised setting is inferior. Why should this be so? Is there any justification (theoretical or experimental)? I can only find vague asides on this point in a few papers.
- Groxx 16y agoIt's not by any means required to be unsupervised, they just tend to be better suited for it, from what I remember. You can train a deep neural network just like a shallow one, but do you know you're training it in the best way? The unsupervised one might key in on something you're not aware of. They are also more capable of dealing with tons of variables (in this case, ~100 million). If you had 100 million yes/no values that needed to be checked, could you really answer all of them accurately? In some cases, sure, but not many, and what's the correct number of variables for simulating a language, and what do they relate to? An unsupervised network will try to fit them automagically, rather than taking your word for it. Unsupervised learning also poses a lot of intelligence-modeling uses, because we're essentially unsupervised for some of the "hardest" tasks to learn. Especially prior to learning to speak. All there really is for feedback is pain, and that doesn't necessarily lead directly to language. But bravura should be able to answer better / more accurately; hopefully they'll chime in too. I'd like to know if there's more to it than what I know, which is pretty minimal.
- bravura 16y agoIf you look at the video, I try to give some intuition about this. The following explanation is an oversimplification to give you intuition into what is going on: The difficulty in training a deep architecture is that, if you use standard backprop, the gradient signal doesn't flow back to the lowest layers. Instead, the few supervised output units pass back a small gradient signal, and at each layer the gradient signal has increasing noise as it gets passed backwards. So the top layers overfit, and the lower layers don't get tuned effectively. The lower layers essentially fire random noise, and the top layers overfit. For this reason, before 2006, no one knew how to train a deep architecture (besides Yann LeCun, with his convolutional architectures, but those were not as general purpose). The breakthrough in 2006 came when Hinton came up with the original DBN algorithm. Bengio et al (http://www.iro.umontreal.ca/~lisa/publications/index.php?page=publication&kind=single&ID=190 http://www.iro.umontreal.ca/~lisa/publications/index.php?pag...) followed up by teasing apart the important steps in training a deep architecture. They were: greedy layerwise training, i.e. construct one layer, then the next layer, then the next layer, etc. unsupervised pretraining When you train a single layer using an unsupervised criterion, the gradient signal is passed backwards through only one hidden layer (the layer you are constructing). And the output layer for an unsupervised criterion has as many units as the input. So the output layer passes back a strong gradient signal, and it doesn't have to travel very far. By doing this unsupervised pretraining in a layer-by-layer manner, the deep network receives a good initialization of its parameters. Then, when finetuning against the supervised criterion using backprop, you can find a better local minimum. The work of Erhan et al show that the effect of unsupervised pretraining is not only regularization, but also improved optimization. Take a look at there work for a great empirical study with large scale experiments and pretty graphs: http://jmlr.csail.mit.edu/proceedings/papers/v9/erhan10a/erhan10a.pdf http://jmlr.csail.mit.edu/proceedings/papers/v9/erhan10a/erh... p.s. feel free to send me an email to followup.