6 ms·
It's a very interesting result, but as always with neural networks we have to keep in mind that what matters is not whether a model can be encoded in a differen
by fchollet 12y ago
It's a very interesting result, but as always with neural networks we have to keep in mind that what matters is not whether a model can be encoded in a different architecture (even at equal entropy), the question is whether the model can be learned in the first place. When you work with shallow nets trained with regular backprop + dropout, you see that their learning capabilities tend to "saturate" much quicker than deep nets. Often with shallow nets, after a point you don't get better results by adding more units or more training data. But deep nets are better able to make use of these extra parameters (extra layers) and extra training data.
Possibly because deep nets conceptually "break down" a learning problem into incremental steps (each new layer being a higher level of representation).
But then again, maybe the problem is simply that we don't have sufficiently good methods for training shallow NNs on large-scale problems. After all, it's only recently that we figured out how to efficiently train deep nets (either pre-training with Autoencoders or RBMs, or through Hessian-free optimization).
- mjw 12y ago> maybe the problem is simply that we don't have sufficiently good methods for training shallow NNs on large-scale problems In a sense, that's what this is though, right? It's a training algorithm for the simpler class of model. It just has to go via the deep model to get there.
- robrenaud 12y agoI like this paper because it turns some of the current intuition about deep nets around. It shows that the current understanding of why deep nets are so good at so many (perceptual) tasks is that the depth buys you a lot. Yoshua Bengio will point out that there are functions that require exponentially more gates to encode when using shallower circuits. This might lead people to believe that deep nets are working so well because they are more fundamentally capable of representing the solutions to problems that people care about in a terse way. But this work proves (at least for this audio task), that there are solutions as good in the solution space spanned by shallow nets with memory usage that we can afford, we just didn't know how to find them.