4 ms·
Really? It's literally been decades, but the last I learned about neural networks, a non-linear activation function was important for Turing completeness.
by mcguire 2y ago
Really?
It's literally been decades, but the last I learned about neural networks, a non-linear activation function was important for Turing completeness.
- kragen 2y agoneural networks aren't turing complete (they're circuits, not state machines) and relu is not just nonlinear but in fact not even differentiable at zero
- jszymborski 2y agoReLU are indeed non-linear, despite the confusing name. The nonlinearity (plus at least two layers) are required to solve nonlinear problems (like the famous XOR example).
- PeterisP 2y agoYes, some non-linearity is important - not for Turing completeness, but because without it the consecutive layers effectively implement a single linear transformation of the same size and you're just doing useless computation. However, the "decision point" of the ReLU (and it's everywhere-differentiable friends like leaky ReLU or ELU) provides a sufficient non-linearity - in essence, just as a sigmoid effectively results in a yes/no chooser with some stuff in the middle for training purposes, so does the ReLU "elbow point". Sigmoid has a problem of 'vanishing gradients' in deep networks, as the sigmoid gradients of 0 - 0.25 in standard backpropagation means that a 'far away' layer will have tiny, useless gradients if there's a hundred sigmoid layers in between.