3 ms·
In very high dimensional spaces (like trying to optimize a neural network with billions of parameters), to be "in a valley", you must be in a valley with respec
by Imnimo 2y ago
In very high dimensional spaces (like trying to optimize a neural network with billions of parameters), to be "in a valley", you must be in a valley with respect to every one of the billions of dimensions. It turns out that in many practical settings, loss landscapes are pretty well-behaved in this regard, and you can almost always find some direction to continue going downward in that lets you go around the next hill rather than over it.
This 2015 paper has some good examples (although it does sort of sweep some issues under the rug): https://arxiv.org/pdf/1412.6544 https://arxiv.org/pdf/1412.6544
- jonathan_landy 2y agoIs the claim that there aren’t local many local minima for high dimensional problems eg in neural network loss functions?
- Imnimo 2y agoYes. To be more specific, it's that nearly all points where the derivative is zero are saddle points rather than minima. Note that some portion of this nice behavior seems to be due to design choices in modern architectures, like residual connections, rather than being a general fact about all high dimensional problems. https://arxiv.org/pdf/1712.09913 https://arxiv.org/pdf/1712.09913 This paper has some nice visualizations.