4 ms·
We were missing two architecture patterns that were needed to get deeper nets to converge: residual nets [1] which solved gradient propagation, and batch normal
by radq 3y ago
We were missing two architecture patterns that were needed to get deeper nets to converge: residual nets [1] which solved gradient propagation, and batch normalization [2] which solved initialization.
[1] Residual nets (2015): https://arxiv.org/abs/1512.03385 https://arxiv.org/abs/1512.03385
[2] Batch normalization (2015): https://arxiv.org/abs/1502.03167 https://arxiv.org/abs/1502.03167
- hzay 3y agoYes, but the tweet is talking about single layer networks!
- sigmoid10 3y agoAlso quasi-linear activation functions (prevent vanishing gradients), tons of regularisation (e.g convolutions) and more adaptive gradient descent (faster convergence). I've still met people in the early 2010s who tried to make neural networks work using only a few dozen units. Academia is pretty slow. What people also forget is that libraries like pytorch or tensorflow simply didn't exist. I wrote my own neural network stacks complete with backpropagation from scratch in c++ back then.
- bravura 3y agoLeCun et al (1989) had backprop working for digit recognition. LeCun, Bottou, et al (2002) in "Efficient Backprop" described techniques for improving backprop algorithms.
- sigmoid10 3y agoRosenblatt had a working perceptron for classifying images in the 1950s (!). And yet it took 60 years before the theory and compute power had developed enough for all of this to be interesting outside of small, purely academic experiments.
- arketyp 3y agoAlexNet predated that though.