4 ms·
Blog spam pass through to the original: http://sebastianruder.com/optimizing-gradient-descent/ http://sebastianruder.com/optimizing-gradient-descent/ Useful re
by highd 10y ago
Blog spam pass through to the original: http://sebastianruder.com/optimizing-gradient-descent/ http://sebastianruder.com/optimizing-gradient-descent/
Useful reference, 6 months old.
Side note, anybody aware of any implementations of the "learning to learn gradient descent by gradient descent" work [0]? I'd love to boost my training time, but I'm worried implementing it myself will just result in an additional set of hyperparameters to tune.
[0] https://arxiv.org/abs/1606.04474 https://arxiv.org/abs/1606.04474
- zo7 10y agoIs it possible to swap the link on this post? That's kind of silly that they just copy the original post. Understanding different optimizers and why they work is important though, especially since often people will just use them without considering what they're actually doing.