4 ms·
> Despite the name 'neuronal networks' that we have history to thank for, it's better think of them as just long compositions of functions with a lot of paramet
by mannigfaltig 9y ago
> Despite the name 'neuronal networks' that we have history to thank for, it's better think of them as just long compositions of functions with a lot of parameters one can tweak.
Actually, the name is not so misleading because `ReLU(Wx + b)` is an approximation of the response of a large ensemble of integrate-and-fire neurons in terms of the firing rate [1] (it models populations of pyramidal neurons in the cortex). The ReLU activation function basically approximates a logarithmic activation function [2]. Obviously many kinds of computations are left out in this model, for example in pyramidal neurons, the integration of the incoming spikes is often sublinear rather than linear; the branches of the neurons that receive signals (dendrites) can themselves perform complex computations such as spatial clustering and logical operators basically sending small (dendritic) spikes forward and backward within an neuron [3]; short-term plasticity is left out entirely and so are countless types of interneurons which inhibit other neurons nearby and measure mean activity across populations of neurons etc.
[1] https://web.stanford.edu/group/brainsinsilicon/documents/Wong_Wang_2006.pdf https://web.stanford.edu/group/brainsinsilicon/documents/Won...
[2] https://www.utc.fr/~bordesan/dokuwiki/_media/en/glorot10nipsworkshop.pdf https://www.utc.fr/~bordesan/dokuwiki/_media/en/glorot10nips...
[3] https://neurophysics.ucsd.edu/courses/physics_171/annurev.neuro.28.061604.135703.pdf https://neurophysics.ucsd.edu/courses/physics_171/annurev.ne...
- eru 9y agoThanks! Yes, you can find that the ReLU does model some aspects---but if memory serves right people didn't switch from sigmoid to ReLU because they were interested in modelling the brain better, but mostly because they were looking for a structure that's simpler to train. Sigmoids notoriously smear out all your backpropagation 'juice' after only a few layers, don't they? (I only mention ReLU because they are definitely simpler in mathematical structure than ye olde sigmoid. Max out is another interesting activation function that's first and foremost inspired by the math / pragmatics of drop-out training, and were any relations to biology more likely to be subsequent discoveries and not intentional design features.)