3 ms·
In my experience tinkering with plain vanilla neural networks, training quality is quite sensitive to the width of the first couple of base layers. In particula
by Scene_Cast2 4y ago
In my experience tinkering with plain vanilla neural networks, training quality is quite sensitive to the width of the first couple of base layers. In particular, using a rhombus / "diamond" shape works best.
Here's my intuition. The output of the first layer is mostly a linear matrix transform. These linearities form the basis for the more complex features. If the network doesn't have enough of these base linear features, it can't compose the higher-level (and lower-dimensional) features later down the line.
- duped 4y agoYou might want to just make the first layer a DCT of the input vector which will split them into decorrelated components.
- wardedVibe 4y agoI see we are re-deriving conv nets
- duped 4y agoNot really, the first principle of not-too-bad ANNs is to use some kind of preprocessed, hopefully decorrellated feature vectors as the inputs. Conceptually that is way easier to implement with intuition and know-how over blindly training an arbitrary structure to do what you want. Ideally the training just derives non-intuitive relationships between these decorrelated input features to give you something useful.
- ausbah 4y agoby rhombus or diamond shaped do you mean the first layer being wider than the input layer - with each following hidden layer being smaller than the previous hidden layer?
- Scene_Cast2 4y agoBy rhombus, I meant something like input -> 20 -> 50 -> 200 -> 50 -> 20 -> output.
- raihansaputra 4y agoI'm still learning the basics of ML, but I assume the advantage of rhombus instead of just having 200*5 layers is training time? Would it have other advantages?