3 ms·
If these claims are true (specifically, that every local minimum is a global minimum), then why did the earlier neural networks have poor performance? Why did w
by rudyl313 10y ago
If these claims are true (specifically, that every local minimum is a global minimum), then why did the earlier neural networks have poor performance? Why did we need advancements like pretraining via stacked RBMs and dropout in order to make deep learning converge on usable/better models?
- argonaut 10y agoA quadratic function is a convex function with a global minimum. Doesn't mean it's a good model for much.
- rudyl313 10y agoBut the tricks/advancements I mentioned are not changing the function. They changed the initial weights and how the cost function was explored.
- aab0 10y agoBut they do change the function. You use RELU now, instead of pretraining.
- rudyl313 10y agoUsing ReLU units is a newer advancement and I agree that changing the activation function does change the cost function. However, before Hinton got all excited about ReLU units, he was still showing huge improvements just by using pretraining and later by using dropout, which shouldn't change the cost function.
- argonaut 10y agoDropout helps with convergence/optimization, sure. The existence of a global minimum says nothing about the time required to reach it. Important to note that dropout isn't as common anymore; it's not a huge win.
- andreyk 10y agoHinton nicely summarized why neural nets used to not work - https://www.youtube.com/watch?v=IcOMKXAw5VA&feature=youtu.be&t=21m29s https://www.youtube.com/watch?v=IcOMKXAw5VA&feature=youtu.be... Quoted: 1. Our labeled datasets were thousands of times too small. 2. Our computers were millions of times too slow. 3. We initialized the weights in a stupid way. 4. We used the wrong type of non-linearity. The pre-training helped with initialization, but later it turned out that just initializing the weights with correct scales for each layer (to deal with dissapearing/exploding gradient effect) worked almost as well with enough data.
- ogrisel 10y agoPoor performance is generally evaluated by considering the validation error. This paper only cares about the training error: it is about the presence or absence of bad local minima in the optimization problem that one has to solve when training a MLP with ReLU activations. The learning and the optimization problems are related but they are not the same :) Dropout is a tool to prevent overfitting (large gap between training error and validation error). This paper does not say anything with overfitting or generalization or the impact of regularization. Also note that unsupervised pre-training via stacked RBMs has proven mostly useless for MLPs with ReLU activations if the network is wide enough and the number of samples in the training set big enough. It is unclear that initialization via unsupervised pre-training can improve the training error or not. I think it mostly has an impact on the validation error (although I am not sure). Furthermore, practitioners tend to stop training before the full convergence on the training set. Instead one generally stops when validation error stops decreasing significantly (early stopping) and one does not really care about the final value of the training loss one could have reached if we had continued training forever. Traditional SGD has a convergence rate that is too slow and in practice it prevents checking whether we are converging to a bad local minima or not on non-toy problems. To sum up: better understanding of the optimization problem is very helpful (in particular to tackle underfitting and reduce training times) but that alone will not ensure that we can build model that generalize correctly to unseen data.