3 ms·
ReLU is very simple in this regard. In the plain form it's just affine transformation followed by 'a viewport'. The mapping trough multiple layers is alternatin
by MAXPOOL 8y ago
ReLU is very simple in this regard. In the plain form it's just affine transformation followed by 'a viewport'. The mapping trough multiple layers is alternating affine transformations and windows into data. Learning is combination of squeezing and rotating data so that it can be seen or unseen from the window and rotating window frames to do the same.
Any extra stuff, like batch normalization between the layers can again introduce more complex nonlinearity.
- jesuslop 8y agoReally agree with the extra simplicity in the ReLU case. Liwen Zhang, Gregory Naitzat and Lek-Heng Lim showed last year that "the family of such neural networks is equivalent to the family of tropical rational maps", where rational functions are quotient of polynomials and "tropical" is in the context of Tropical Geometry, where instead of the typical "plus, times" ring one uses a "min, plus" ring that has somewhat unexpected applications. For instance, with its module theory one can calculate minimum costs paths in graphs just as one does for reachability starting with the adjacency matrix of a graph over a boolean semiring. arXiv:1805.07091v1