5 ms·
Great questions. Happy to answer them here. First of all, this work builds on Hinton et al.’s second paper, the one about EM routing of matrix capsules, from
by fheinsen 7y ago
Great questions. Happy to answer them here.
First of all, this work builds on Hinton et al.’s second paper, the one about EM routing of matrix capsules, from last year: https://ai.google/research/pubs/pub46653 https://ai.google/research/pubs/pub46653 This work is only minimally related to his previous paper (Sabour et al.'s paper) from two years ago!
RESPONSES TO #1:
* The same algorithm also achieves SOTA in another domain, natural language. Same code. I think it’s significant that the same code, without change, produces SOTA in two domains. See the README and tables 3 and 4 in the draft paper. Don't you think this is significant?
* It requires fewer parameters: 272K instead of 310K for Hinton et al. (2018)’s model and 2.7M for the best performing CNN on record (Cireşan et al.); see table 2. That’s 10x fewer parameters than the best performing CNN on record.
* It requires an order of magnitude less training: 50 epochs instead of 300 for Hinton et al. (2018)'s model.
* It’s trained with minimal data augmentation, unlike Hinton et al.’s and Cireşan et al.’s models (the latter, in particular, uses a ton of data augmentation). Also, unlike Hinton’s model, it accepts full-size images instead of 32x32 crops that are 9 times smaller. Finally, we do not measure accuracy as a mean of multiple crops. So, the model has fewer parameters, requires less training, and has greater capacity.
* It seems to be learning a form of "reverse graphics" on its own, from only pixels and labels, without having to optimize explicitly for it. See the README, figure 4, and the 24 plots and captions in supplemental figures 6 and 7. This is rather significant, don't you think?
RESPONSES TO #2:
* As far as I know, the best attempt at recreating Hinton et al.’s work on EM routing is by Ashley Gritzman at IBM, in July of this year -- only a bit over two months ago. As far as I can tell, his model does not come close to matching Hinton’s performance:
https://arxiv.org/abs/1907.00652 https://arxiv.org/abs/1907.00652
https://github.com/IBM/matrix-capsules-with-em-routing https://github.com/IBM/matrix-capsules-with-em-routing
https://medium.com/@ashleygritzman/available-now-open-source-implementation-of-hintons-matrix-capsules-with-em-routing-e5601825ee2a https://medium.com/@ashleygritzman/available-now-open-source...
* There have been a few other efforts, all of which seem to fall short of Hinton's performance. Gritzman does a good job of covering those other efforts in his Medium article. None of these efforts propose any new ideas, as far as I can tell.
RESPONSES TO #3:
* Me too. So does Hinton: https://openreview.net/forum?id=HJWLfGWRb https://openreview.net/forum?id=HJWLfGWRb ... and so does everyone else.
* Alas, as Paul Barham and Michal Isard at Google Brain showed earlier this year, currently it can be challenging to scale capsule networks to large datasets and output spaces, in some circumstances, due in part to current software (e.g., PyTorch, TensorFlow) and hardware (e.g., GPUs, TPUs) systems, which are highly optimized for a fairly small set of computational kernels, in a way that is tightly coupled with memory hardware, leading to poor performance on non-standard workloads, including basic operations on capsules. Source: Barham and Isard (2019) - https://dl.acm.org/citation.cfm?id=3321441 https://dl.acm.org/citation.cfm?id=3321441 (the PDF is available for free download at that link).
* My draft paper mentions Barham and Isard’s work.
- p1esk 7y ago1. Both Hinton’s capsules papers have been released at the same time (Oct 2017). You can see the first comment on OpenReview page for the EM paper is dated Nov 2017. From what I remember, the two papers appear very similar with the main difference in how the routing is implemented. 2. You cite a convnet result from 2011 (!). Don’t you think a modern convnet would do vastly better on this task? 3. Could input size play a role? Did you try feeding 96x96 inputs to the models you’re comparing against, to see if they also benefit from it? 4. I’m a bit confused as to why other implementations failed to reproduce Hinton’s results given that he open sourced their code (link in the first OpenReview comment). 5. Ok, Imagenet is too slow, how about Cifar-10? What would it take to reach, say, 95%? That would be equivalent to a well trained Resnet-18. If you can show such result, I personally would become more interested, because I worked quite a bit with Cifar-10, but not with Norb. I think you might be onto something, but it’s still not clear that capsules approach is scalable and ultimately superior to plain convnets.
- fheinsen 7y agoI’m surprised you did not comment on the fact that my version of EM routing also achieves SOTA on another domain, natural language. Same code. Here are the answers to your questions: 1. The final, published version is stamped “ICLR 2018,” so I used that year. 2. I don’t know if a conventional CNN can do this with 10x fewer parameters, while also learning to do a form of “reverse graphics” without explicitly optimizing for it. (I wouldn’t know how to get a CNN to do that without explicitly making it a training objective.) 3. IIRC, the convnet model from 2011 accepts 96x96 images. As to why Hinton et al. downsample images to 9x smaller, I suspect (but don’t know for sure) they had no choice to conserve memory and computation using their version of EM routing. I was able to reduce memory and computation with my variant of EM routing (by between one and two orders of magnitude) by setting the first routing layer to accept a variable number of inputs, without regard to location in image. 4. Me too. But you asked me about work other than Hinton’s, and that’s all I could find! 5. CIFAR10 is on the to-do list (work permitting!) :-)
- p1esk 7y agoHow does a regular convnet do on another domain? Learning to do “reverse graphics” is only useful if you can show it is the reason behind performance improvement, compared to a plain convnet. Until we have cifar-10 results it’s not clear. What I’m saying is - no one has yet demonstrated a clear superiority of any capsules based model to the best available plain convnet. Even on cifar-10. Looking forward to your results!