5 ms·
Unfortunately, the 1-step optimal learning rate often differs massively from the long-horizon best choice: https://arxiv.org/abs/1803.02021 https://arxiv.org/a
by Straw 7y ago
Unfortunately, the 1-step optimal learning rate often differs massively from the long-horizon best choice:
https://arxiv.org/abs/1803.02021 https://arxiv.org/abs/1803.02021
Due to the fact that larger LRs can result in worse immediate performance but better progress in low signal-to-noise ratio directions.
- 0-_-0 7y agoI suspect that's because higher LR leads to better exploration of the opt. surface, i.e. it works as an implicit regularizer. The ideal solution would be to develop better regularizers to go with the better optimizers, instead of relying on the noise in the worse optimizer for implicit regularization.
- Straw 7y agoI completely agree, we shouldn't be depending on our optimizers do some approximate Bayesian inference- an optimizer should optimize only. However, I think it's a different effect- even purely in terms of optimizing the training loss, on a quadratic (with noisy gradients), the short-horizon bias effect exists.