4 ms·
I don't think we can say for sure that early stopping is the main reason deep networks generalize. Double descent [1] shows that models continue to improve even
by maxwells-daemon 6y ago
I don't think we can say for sure that early stopping is the main reason deep networks generalize. Double descent [1] shows that models continue to improve even once they've "interpolated" the training data (fit every point perfectly), and critical periods [2] suggest that the early part of training is responsible for most of the generalization performance even though much of the numerical improvement happens later.
Overall it looks like gradient descent is a strong regularizer -- we know it tends to prefer small and low-variance weights, for example. So part of deep generalization has to do with how SGD is able to pick "good" features early, and then optimization pushes the unimportant weights to zero later (hence lottery tickets).
[1] https://openai.com/blog/deep-double-descent/ https://openai.com/blog/deep-double-descent/ and other papers.
[2] https://arxiv.org/abs/1711.08856 https://arxiv.org/abs/1711.08856 and others.
- owenshen24 6y agoTo add some more context, here's a rather readable summary: https://www.greaterwrong.com/posts/FRv7ryoqtvSuqBxuT/understanding-deep-double-descent https://www.greaterwrong.com/posts/FRv7ryoqtvSuqBxuT/underst...
- machinelearning 6y agoI think you misunderstood the point of deep double descent. The x-axis is not number of training epochs, it is model capacity. I think you'd be interested in https://arxiv.org/abs/1611.03530 https://arxiv.org/abs/1611.03530. It discusses how SGD is an implicit regularizer. We also actually want high variance weights for symmetry breaking.