4 ms·
Hi, I am a research engineer in Yann LeCun's group at Facebook. I hate to seem to be picking a fight, but you're a prominent poster here, and this comment seems
by kmavm 12y ago
Hi, I am a research engineer in Yann LeCun's group at Facebook. I hate to seem to be picking a fight, but you're a prominent poster here, and this comment seems likely to garner a fair amount of attention. Unfortunately almost every sentence you've written betrays a subtle misunderstanding of the space, and the totality is quite misleading.
> The naive approach (start at a "random" weight setting, use gradient descent) that works on a convex error surface fails catastrophically on deep, complex neural nets. (Shallow neural net training is a non-convex problem as well, but seems to be "less non-convex" in practice.)
Starting at a random (no scare quotes needed) point in weight space and SGD'ing is exactly what Alex Krizhevsky, and all of the derivative convnets over the last two years, did and do. It works just fine; I sit at work doing it all day long. You need to have enough data to train on, big enough models, and enough flops to train the big models on the big data before your interns' grandchildren die. We have all of the above now. Aside: even single-layer neural networks do not have convex error surfaces; convexity, and funky error surface geometry, is not a relevant distinction between shallow and deep nets. There have been no magical optimization breakthroughs, it's still SGD with the same herbs and spices that were used in the 90's (momentum, e.g.).
> The problem is that convergence, for single-layer nets, can be very slow (especially given that you're often doing stochastic gradient descent when working large data sets).
"Stochastic" gradient descent just means doing lots of weight updates per epoch. Ceteris paribus, training on large, redundant data sets, stochastic converges faster than batch because it gets to consider more points in the weight space than batch per pass over the data. The problem with single-layer neural nets is not that they converge slowly; the problem is that the layer size needs to grow exponentially with the task size. Single-layer neural nets' universal approximation power is thus not of great practical consequence. The power of deep nets is the power of composition: f(g(h(x))) is a strictly more powerful model than f(x) holding the number of parameters constant.
> Even now, making deep neural nets not sensitive to initial starting conditions is an unsolved problem, but there's been a lot of progress.
You just initialize with a Gaussian ball around zero and explore whatever valley in the error surface you happen to be in. Works 100% dandy.
> I would hazard the guess that the convolutional technique is a lot more useful in deep neural networks than it is in single-hidden-layer neural nets.
It doesn't really make sense to talk about a "single layer convolutional net", because if you only have a single layer, and all you can do is convolve with it, then the output of your net will necessarily be a big pile of filtered versions of the input image. Unless your task is specifically to learn a target set of filters, it would make no sense to have a single layer convnet.
- robrenaud 12y agoHave you read Rich Cuarana's work on approximating the behavior of deep nets with single hidden layer networks? http://arxiv.org/pdf/1312.6184.pdf http://arxiv.org/pdf/1312.6184.pdf From that paper, it seems like it's just that deep nets are easier to train to find state of the art solutions than shallow nets. It's not that shallow nets are inherently incapable of performing as well. Shallow nets can mimic/approximate a given well trained deep net and preserve almost all of the accuracy. So it's not the case that a good solution to the task doesn't exist in the set of hypothesizes spanned by sized shallow net, it's just that people don't know how to effectively find the right parameters for the shallow nets.
- srean 12y agoMinor comment, please link the abstract and not the pdf. Those who want to follow up can pull the link to the pdf from the arxiv abstract.
- kastnerkyle 12y agoSo far this is only shown on TIMIT and MNIST, which are pretty trivial datasets, so it may be dataset dependent. One of the authors is giving a talk at INRIA in France this month, and the abstract mentioned CIFAR10 (possibly unpublished results). If they have manged to compress a CIFAR10 network, that is a strong indicator that ImageNet networks could be compressed in the same way... but no one has published any results in this regard to my knowledge. However this is an active area of research for me, and I hope to explore it more soon. I think there are better ways to approximate than this, but this work at least shows that it may be possible.
- srean 12y ago> Aside: even single-layer neural networks do not have convex error surfaces; Single layers can indeed be made to have convex error surfaces fairly easily. One can do so by matching the error/loss function with the squashing/link function. What some old NN folks got wrong was mixing up square loss with logistic function, that is an unhealthy mix. Now if one were to use KL divergence instead of square loss then one would indeed have a convex loss function. In fact this would be nothing but logistic regression. One can however push this idea further, with any choice of a monotonic squashing function one can derive a 'matching' loss that would give you a convex loss. Classical statisticians know this and call it with a different name: canonical generalized linear models. I am not from that tribe, mine is more ML we may perhaps call it minimizing Bregman loss. Just so that its clear I am talking about single layer networks not single hidden layer networks, there are plenty of cases were the former is useful. > There have been no magical optimization breakthroughs It is arguable whether Hessian free methods, contrastive divergence or auto-encoder based training methods qualify as 'breakthroughs' but they have definitely equipped invigorated researchers in this broad area with their capabilities.