4 ms·
The reasons to use ReLUs are sparsity and improved gradient flow. ReLUs encourage sparsity because when the input to the ReLU is less than 0 as the activation b
by jimfleming 10y ago
The reasons to use ReLUs are sparsity and improved gradient flow. ReLUs encourage sparsity because when the input to the ReLU is less than 0 as the activation becomes 0. This means some fraction of activations in a given layer will be omitted which can encourage better representations. They also have improved gradient flow because the gradients are zero or constant and thus don't suffer from vanishing/exploding gradients.
In deep learning, I would _generally_ not look towards biology for the reasons behind why things are done as this is usually an after-the-fact explanation. When in doubt, blame the gradients.
- ma2rten 10y agoLeakyReLUs work as well or better than ReLUs, so it can't be because of sparsity.
- ekelsen 10y agoLeakyReLUs often have very small slopes on the negative side; this can help solve the problem of no gradients getting through a layer because the activations were all 0. But I would say it still has a sparsity effect on a trained network because of how small the slope usually is compared to one.
- argonaut 10y agoThat is not sparsity. In machine learning there is a very strong distinction between values that are exactly 0, and values that are close to zero (see: difference between L1 and L2 regularization).
- ekelsen 10y agoL1 regularization does not lead to values that are _exactly_ 0 either.
- makeset 10y agoYes, it does. You might have been thrown off by the fact that L1 regularization is not L0 regularization, i.e. it doesn't explicitly limit the number of nonzero coefficients. Still, the linearity of L1 constraint boundaries creates spikes in directions with zero components, thus forcing constrained solutions to occur where many variables are driven to exactly zero. See here: https://en.wikipedia.org/wiki/Lasso_(statistics)#Geometric_interpretation https://en.wikipedia.org/wiki/Lasso_(statistics)#Geometric_i...
- ekelsen 10y agoIf we have a parameter x, and some cost function J(x), then with L1 regularization the cost function would be J(x) + beta * abs(x). The derivative of that loss with respect to x would be J'(x) + beta * sgn(x). So using some variant of SGD (which is what basically all neural network training does these days) we would essentially update x as: x = x - alpha * (J'(x) + beta). (The specifics depend on the algorithm, but it doesn't change the result). So for x to end up as _exactly_ 0, we have to be extremely lucky, which in practice I have never observed. Using L1 regularization definitely leads to small weights, but not to ones that are _exactly_ 0.
- argonaut 10y agoYour conclusion is theoretically false. You can prove that L1 regularization is equivalent to taking the optimal unregularized parameters, setting to parameters below a threshold to 0 (the threshold depends on the regularization parameter), and penalizing the other parameters.
- ekelsen 10y agoYes, but how do you actually optimize that loss in practice? I'm not saying that a perfect solution with an L1 penalty wouldn't have weights exactly equal to 0. I'm saying that with the optimization techniques that are commonly used, you don't end up with exact zeros.
- jimfleming 10y agoThe sparsity encouraged by ReLUs is one reason they are used but not to the exclusion of all other activations. ReLU variants can indeed outperform standard ReLU[0] but sometimes sparsity is a desired property. For example, in generative models. [0] https://arxiv.org/abs/1505.00853 https://arxiv.org/abs/1505.00853