4 ms·
Its quite important that its not so much about "non-convexity", its more about "random initialization". All these results, and there are others, more sophistica
by cosmic_ape 8y ago
Its quite important that its not so much about "non-convexity", its more about "random initialization". All these results, and there are others, more sophisticated, basically say that after random init everything is already good, and the gradient descent just has to train the highest layer, without screwing up the previous ones. They are all pretty naive that way.