3 ms·
I think this is a very good point. Years ago people were worried about gradient descent getting stuck at a local minima, plausibly because this problem is very
by kahoon 9y ago
I think this is a very good point. Years ago people were worried about gradient descent getting stuck at a local minima, plausibly because this problem is very obvious in a 3 dimensional space. In higher dimensions however this problem seems to go away more or less and a lot of worrying about the issue seems to be the result of lower dimensional intuitions wrongly extrapolated to higher dimensions.
- Eliezer 9y agoStuff got stuck in local minima for years before we learned about stuff like momentum and dropout and dropped a ton of GPU power on it.
- llamaz 9y agoWhen I was implementing a neural network for a university assignment (2 years ago so my memory might fail me), we had to run our algorithm multiple time with different starting positions, then take the minimum of those local minima. I'm not sure what momentum and dropout are, but I agree with Eleizer, without these things (which I didn't use) local minima are a problem.
- eat_veggies 9y agoDropout is where you randomly remove neurons from your network during training, which prevents them from depending too much on specific neurons (making the output more generalizable). It was developed in 2014 so it would have been brand new tech back when you were in your class.
- alexcnwy 9y agoI think you're misinterpreting the parent who is saying that local minima are not a problem in high dimensions because there is always a dimension to move in that reduces the loss (unlike in lower dimensions where you can get stuck in a point across all dimensions that cannot be locally improved upon)
- llamaz 9y agoI still don't understand what the parent is talking about then. Could you please restate the explanation using math notation/terminology?
- bobby_the_whale 9y agoCan you write down your statement in a formal language such that we can prove the negative to your statement such that you might get the idea to stop talking about things you do not understand?
- david-gpu 9y agoNot the person you replied to, but your comment was both rude and incorrect enough that I feel the need to reply. See for example http://www.offconvex.org/2016/03/22/saddlepoints/ http://www.offconvex.org/2016/03/22/saddlepoints/ for some discussion on this.
- bobby_the_whale 9y agoFunny you say that. I have the impression you can't even write down what correctness means. If you have to refer to an external party that is also lacking in rigor, please don't. EDIT: Have you even read that article yourself and in particular the NP-Hard bit? That directly contradicts the idea that you can escape from local minima. The only thing you can hope for is escaping from some local minima or that your problem actually was easy to begin with. Computers have never in their history solved hard problems for non-trivial problem sizes, they have merely approximated them. Neural networks have been used to solve problems of practical interest, but any extension of that claim just makes you look like a clown.
- bovine3dom 9y agoSay you have N random walks. The probability that the second derivative at any point is of the same sign for all walks decreases with N. Right?
- bobby_the_whale 9y agoYou know the expression "not even wrong"? That's exactly what this is. If you take the lim_{N->\inf} it's true, sure. Except that's a trivial result. We already have a lower existing upper bound available for solving neural network optimization problems.