4 ms·
My NN knowledge is minimal, but from what I recall, the claim always was that you don't need more than 1 hidden layer; and too many hidden neurons results in ov
by ajays 14y ago
My NN knowledge is minimal, but from what I recall, the claim always was that you don't need more than 1 hidden layer; and too many hidden neurons results in over-fitting.
What changed?
- tgflynn 14y agoThe claim that you don't need more than 1 hidden layer is based on a mathematical theorem that says roughly that a sufficiently large 3 layer net exists that can fit any sufficiently smooth function. One has to look in detail at things like "sufficiently smooth" and "sufficiently large" when applying mathematical theorems to real problems - and that's a step practitioners often seem to neglect when looking for rules of thumb. Also just because a net exists doesn't mean that any given training algorithm is likely to find it. As for overfitting the best way to reduce it is to use more training data and I believe the nets discussed in these papers have been trained on some of the largest training sets ever used. The other problem with deep networks is that they have been considered very difficult to train with backpropagation due to vanishing or exploding gradients. I think the recent major algorithmic developments have mainly involved methods to mitigate these problems.