8 ms·
Just started watching the video so they may mention this but one thing I find fascinating is that some recent work suggests that the optimization algorithm (usu
by micro_cam 9y ago
Just started watching the video so they may mention this but one thing I find fascinating is that some recent work suggests that the optimization algorithm (usually stochastic gradient descent) or complexity of the loss surface (ie having lots of local minima that are almost as good as the global maxima) may actually be seen to induce a kind of regularization prior.
Ie these seemingly really complex models are actually biased to find simpler solutions that generalize well in a way that turns out to often work better then trying to explicitly learn a simple model.