4 ms·
I'm going to check back later to see if anyone manages to reproduce it. Perhaps by the time it's presented at NIPS. A twitter conversation reflecting some scep
by imurray 10y ago
I'm going to check back later to see if anyone manages to reproduce it. Perhaps by the time it's presented at NIPS.
A twitter conversation reflecting some scepticism, but agreeing it would be interesting if it all checks out:
https://twitter.com/fchollet/status/771862837819867136 https://twitter.com/fchollet/status/771862837819867136
- cs702 10y agoI'm less skeptical than fchollet (creator of Keras, for those here who don't know), but agree that we need to wait until the usual suspects at Google, FaceBook, Toronto, Montreal, Stanford, etc. have replicated this. In all likelihood the team will release code soon, either before or after NIPS, so we will all be able to check things out for ourselves.
- cs702 10y agoOne of the authors, Zhangyang Wang, just wrote this on his personal page: "We have discussed and decided to work on a software package release, perhaps accompanying it with a more detailed technical report in the future. Once the software package is ready, we will update everybody." http://www.atlaswang.com/ http://www.atlaswang.com/
- joe_the_user 10y agoPaper withdrawn https://arxiv.org/abs/1608.04062 https://arxiv.org/abs/1608.04062 It's kind of an odd thing. I (random non-academic amateur) actually spent a bunch of time trying to parse the paper, which was kind of a combination of interesting ideas and incomprehensible ambiguities. One real academic researcher also put some time into it. The good part of the paper is explained here. My guess is the problem is going from ARM to SARM. http://gabgoh.github.io/SARG/ http://gabgoh.github.io/SARG/ While I'm sure most people involved think of the experience as a wash, I feel like I learned a bunch about deep learning in the process. PS, also sad that the author did this.
- gabrielgoh 10y agoHi, I'm the author of the blog post. I added a blurb to the beginning the blog post explaining all the drama, and precisely what claim was made that was withdrawn. The problem is not in the ARG->SARG approximation, but the bit on unsupervised pretraining. The paper could stand on its own without that section, but without that result it would have been a significantly more mediocre NIPS submission. Hope this clarifies things.
- nl 10y agoYour blog post was excellent, btw.
- joe_the_user 10y agoFirst, thanks for the excellent blog, it gave me a better idea what was happening As far as the ARG-> transformation goes, maybe that's just something I don't get, I can see how one goes from sparse encoding to repeated ARG-type transformations and how this repeated application approximates the solution of a sparse encoding problem. And it is suggestive that these application look like a layers of a neural net. But when you switch to stacking, what are you doing? Solving one sparse encoding problem then another? What analogy is there to say this works ... or that it would work better than just single sparse encoding? At that point, is it just "try it and see?" One of the impressions I got from scanning the literature is that deep nets are kind of generally treacherous beasts - just getting a locally 1st layer may not be desirable off the bat. People have settled on backpropagation for very subtle reasons. See "Overfitting in Neural Nets: Backpropagation, Conjugate Gradient, and Early Stopping", Caruana, Lawrence, et. al where backpropagation finds better solutions than the "more powerful" conjugate gradient method.
- gabrielgoh 10y agoyou are right. you are using the output of the previous sparse solution as input into the new one, i.e. stacking sparse coders. Your second question of why this is a good idea is the million dollar question. Its pretty much "lets try it and see", with some heuristic reasoning thrown into the mix (its mirrors the brain, it abstracts information, etc, etc). btw, I don't think people use early stopping anymore. It's been replaced by more powerful forms of regularization, such as dropout. The deep learning world is getting more tame, and that makes me happy.