5 ms·
The “physicist” explanation that I heard (meaning non-rigorous but good for building intuition) is that at every point where the derivative vanishes, for suitab
by zachf 6y ago
The “physicist” explanation that I heard (meaning non-rigorous but good for building intuition) is that at every point where the derivative vanishes, for suitably random functions, every direction you move in will either be a direction where you increase or decrease at about 50-50 odds. In D dimensions there are 2D independent directions to rise or fall in (e.g. in D=2 dimensions, there’s north south east west), so there’s about 1/2^(2D) odds that any given critical point is actually a minimum if those probabilities are independent. That gets small really fast at large D.
Obviously this is not rigorous, though.
- maxlamb 6y agoThat makes sense and explains why it's unlikely to have local minimas, but I don't get the "no minimal at all" argument. Why no global minimas at all, just because of high dimensionality?
- blackbear_ 6y ago> but I don't get the "no minimal at all" argument. That's just wrong, no need to think about it. Only unbounded loss functions don't have minima, but using such a loss function would not make sense.
- moultano 6y agoThis isn't true. A loss function that asymptotes can also have no minima, which commonly used loss functions do.
- bananaface 6y agoDoesn't the fact the network is discrete (floats have a maximum precision) mean this isn't actually the case? There's a finite number of states the net can be in, and one (or more) is best.
- dougabug 6y agoIt’s a heuristic argument that critical points are extremely unlikely to be local minima (ie positive definite second derivative). Loss surfaces of DNNs do typically have a global minimum (zero if they fit the training data exactly).
- PeterisP 6y agoArguably, a DNN seems likely to have many global minimums - given the level of (over)parametrization commonly used, a set of parameters that gets the lowest possible loss won't be unique, there will be huge sets of parameters that give exactly identical results.
- dougabug 6y agoDue to symmetry, at least, there are many global minima, but with the same minimum value.
- an_d_rew 6y agoBut you have to be careful about that word "independent". There's a reason that things like 3D protein structure estimation, for example, are still very difficult problems, because none of the coordinates are even approximately independent of the others. So you're back to a standard "minimization is really difficult" even in ultra-high dimensional spaces.
- moultano 6y agoYeah I was thinking about that as I was writing and trying to convey why I feel like deep models are different. I think one way of thinking about it is that protein structure, even though it has lots of parameters, it is all happening within the confines of 3D space. A protein that could move in lots of dimensions at once, could probably reliably fold much more easily, and it would be easy to find this structure.