4 ms·
And for anyone who wants to know why unmodified gradient descent may be considered a piece of shit in certain circumstances http://wikipedia.org/wiki/Rosenbroc
by hnuser355 8y ago
And for anyone who wants to know why unmodified gradient descent may be considered a piece of shit in certain circumstances
http://wikipedia.org/wiki/Rosenbrock_function http://wikipedia.org/wiki/Rosenbrock_function
Gradient descent with a good line search (Wolfe conditions) applies to the multidimensional case should converge to min, but it might take you thousands of iterations. Newton’s method or something might take <50.
But machine learning practitioners will know why gradient algorithms are often preferred despite this
- repsilat 8y agoI'm not in ML, but one thing I remember from school (a million years ago, before ML was big) is that Newton's method needs a matrix inversion every iteration, which is expensive when you have a lot of decision variables. Not sure why other old school algorithms like conjugate gradient aren't used though. I guess for rectified linear activation functions the second derivative isn't useful. Maybe that's it.
- marcosdumay 8y agoIt's hard to propagate Newton's method over layers on a neural network.
- stochastic_monk 8y agoThere are also a lot of conditions required for Newton’s method to work which you don’t have with neural networks.
- marcosdumay 8y agoIs there some condition that makes the method not work at all? I could never find a showstopper (granted that I have thought about this for a few hours when first studying the subject), only stuff that slowed it down so gradient descent became better (and honestly, I am still not sure that can not be fixed).