4 ms·
This is not necessarily a trivial fact, but I wouldn't call it incredible. It says a net trained with gradient descent can fit the input perfectly if its much
by cosmic_ape 8y ago
This is not necessarily a trivial fact, but I wouldn't call it incredible. It says a net trained with gradient descent can fit the input perfectly if its much larger than the input.
But this says nothing about generalization, i.e. performance on test set -- which is what we really want.
- iandanforth 8y agoImportantly not SGD, just GD.
- cosmic_ape 8y agoyou're right, fixed.