4 ms·
Multilayer neural networks weren’t really a viable tool until the backpropagation algorithm for determining internal parameters was developed in 1985 (cf https:
by kd5bjo 5y ago
Multilayer neural networks weren’t really a viable tool until the backpropagation algorithm for determining internal parameters was developed in 1985 (cf https://apps.dtic.mil/sti/pdfs/ADA164453.pdf https://apps.dtic.mil/sti/pdfs/ADA164453.pdf )
- _0ffh 5y agoI think you'll find that backpropagation was essentially developed (multiple times) during the 60s in the field of control theory and first implemented in the early 70s. Ed: To be clear, the idea to use them to adapt the weights of NNs was also from the 70s but only rediscovered and applied to MLPs by at least two independent groups/individuals in the 80s.
- mindcrime 5y agoI think you'll find that backpropagation was essentially developed (multiple times) during the 60s in the field of control theory and first implemented in the early 70s. Indeed. The book Talking Nets: An Oral History of Neural Networks[1] covers a lot of this ground. Read it and you'll see many people who were involved in the early history NN's mentioning how backprop was discovered and re-discovered over and over again. [1]: https://www.amazon.com/Talking-Nets-History-Neural-Networks/dp/0262511118 https://www.amazon.com/Talking-Nets-History-Neural-Networks/...
- enchiridion 5y agoWhat really made the difference was non-linear activations. Without a non-linearity depth doesn’t buy you anything.
- visarga 5y agoIt took us long enough to settle on ReLu. Step function activations were brutal.
- canjobear 5y agoI seem to recall that before ReLU it was usually tanh.
- whatshisface 5y agoTanh belongs to a class of functions called "smooth steps" which I guess was being abbreviated as "step." Obviously the derivative of a step is zero everywhere it's defined so backprop wouldn't work.
- l33tman 5y agoAn interesting thing is that ReLu (which is absolutely the most common activation function) has a zero derivative over half the space its defined, and backprop still works (most of the time :). There are leaky ReLU to give some slope to the negative side, but it seems in most normal deep learning networks this is not required, there are always some neurons that aren't in the off regime to "catch the error".
- _0ffh 5y agoOne of a whole class of functions called sigmoidal, usally the sigmoid or the tanh function. These activation functions dominated the field for a very long time.
- l33tman 5y agoInteresting counterpoint: a NN is not just activation. It's also learning, and multilayer stacks learn differently depending on depth even if they are entirely linear. This seems to be an artefact due to the optimizer in use (SGD for example). The jury seems to be out still on why this is.. the issue has been known since long, and I recently saw a short paper Yann LeCun posted about exactly this as well, they stacked a bunch of completely linear layers on top of a "normal" DL stack and got differences in the final system, even though you can collapse all the linear layers at any time to a single linear mixing layer.
- greenbit 5y agoTruth! Without nonlinearity, each layer is essentially just a matrix multiplication. You could just as well use a single equivalent layer that represents the product of the layers' matrices. Nonlinearity lets a single node partition the input space into two (fuzzy) equivalence classes. That's powerful stuff.