4 ms·
Playing around with classic three layer feed-forward classifier networks decades ago, it became apparent that successful networks tended to converge on a sort o
by greenbit 2y ago
Playing around with classic three layer feed-forward classifier networks decades ago, it became apparent that successful networks tended to converge on a sort of standard solution: layer 1 would be various linear discriminators - each node would cut the input space into two half spaces, such that any input point was either in the selected half-space or not, essentially reducing the N dimensional real-valued input to some number of bits. Layer 2 nodes tended toward AND functionality, using some subset of the layer 1 outputs to identify convex regions of the input space. Layer 3 nodes tended to OR together outputs of layer 2 to combine convex regions into arbitrary subsets of the input space. I.e. more often than not, the 2nd and 3rd layer were just doing Boolean logic.
Of course, occasionally there would be a system where this didn't fully happen, and you'd get some node/s doing some funky stuff with the linear parts of the activation function, but that doesn't negate the fact that such networks can and often do find the slice/AND/OR solutions.