3 ms·
The improvements to gradient descent techniques over vanilla gradient descent as well as ideas like batch normalization and residual nets have made even moderat
by highd 10y ago
The improvements to gradient descent techniques over vanilla gradient descent as well as ideas like batch normalization and residual nets have made even moderately large networks that were previously impossible to train now possible, let alone computationally tractable. Most of this work was done in the last 5-6 years. Furthermore, training can now be made drastically more robust, with far fewer empirically set parameters. This has made deep neural nets far more useful in practical application.
- siavosh 10y agoI don't doubt that significant progess has been made in the last five or so years. Meta papers just make me curious given my past experience: the field had ended up with so many tunable parameters that no one understood, it gave rise to whole new papers to 'learn' some arbitrary set of those parameters by a different method. The dirty secret was you were simply adding more parameters. Thus the real issue was the research field had stalled and hit a temporary dead end.