5 ms·
Show HN: Beating Hinton et al.'s capsule net with fewer params and less training
Hello HN,
I recently posted a work-in-progress paper, along with code necessary for replicating all its results, at:
https://github.com/glassroom/heinsen_routing
Among other things, the code in this repo outperforms Hinton et al.'s recent state-of-the-art result in visual recognition[0] while requiring fewer parameters and an order-of-magnitude fewer training epochs.
Most of the original research we do at work tends to be either proprietary in nature or tightly coupled to internal code, so we cannot share it with the world. In this case, however, I was able to remove all traces of internal code and release this as stand-alone open-source software without having to disclose any key IP.
I've reached out to academics in different groups for feedback, and the response so far has been positive, although most have only skimmed the paper. It will likely take a few weeks to get proper feedback from academia.
In the meantime, I figured there are a lot of super-smart, knowledgeable people on HN who would love to take a look at this and share their thoughts. Please feel free to ask questions. Let me know what you think!
[0] https://ai.google/research/pubs/pub46653
- p1esk 7y agoHey, congrats on publishing! 1. Could you briefly summarize your algorithm (novelty, how it’s better, why it’s better, etc)? 2. Since the original paper there have been dozens of published attempts to improve upon it. Do you compare your results to the latest in capsules research? 3. I personally would like to see Imagenet results. Norb a toy dataset. If you beat EfficientNet in terms of both accuracy and number of params/flops, many people will be impressed (including Hinton). Or match the performance of a good convnet using 1/10 training data. Don’t take this the wrong way, but two years after the original paper Norb results, no matter how good, are underwhelming.
- fheinsen 7y agoGreat questions. Happy to answer them here. First of all, this work builds on Hinton et al.’s second paper, the one about EM routing of matrix capsules, from last year: https://ai.google/research/pubs/pub46653 https://ai.google/research/pubs/pub46653 This work is only minimally related to his previous paper (Sabour et al.'s paper) from two years ago! RESPONSES TO #1: * The same algorithm also achieves SOTA in another domain, natural language. Same code. I think it’s significant that the same code, without change, produces SOTA in two domains. See the README and tables 3 and 4 in the draft paper. Don't you think this is significant? * It requires fewer parameters: 272K instead of 310K for Hinton et al. (2018)’s model and 2.7M for the best performing CNN on record (Cireşan et al.); see table 2. That’s 10x fewer parameters than the best performing CNN on record. * It requires an order of magnitude less training: 50 epochs instead of 300 for Hinton et al. (2018)'s model. * It’s trained with minimal data augmentation, unlike Hinton et al.’s and Cireşan et al.’s models (the latter, in particular, uses a ton of data augmentation). Also, unlike Hinton’s model, it accepts full-size images instead of 32x32 crops that are 9 times smaller. Finally, we do not measure accuracy as a mean of multiple crops. So, the model has fewer parameters, requires less training, and has greater capacity. * It seems to be learning a form of "reverse graphics" on its own, from only pixels and labels, without having to optimize explicitly for it. See the README, figure 4, and the 24 plots and captions in supplemental figures 6 and 7. This is rather significant, don't you think? RESPONSES TO #2: * As far as I know, the best attempt at recreating Hinton et al.’s work on EM routing is by Ashley Gritzman at IBM, in July of this year -- only a bit over two months ago. As far as I can tell, his model does not come close to matching Hinton’s performance: https://arxiv.org/abs/1907.00652 https://arxiv.org/abs/1907.00652 https://github.com/IBM/matrix-capsules-with-em-routing https://github.com/IBM/matrix-capsules-with-em-routing https://medium.com/@ashleygritzman/available-now-open-source-implementation-of-hintons-matrix-capsules-with-em-routing-e5601825ee2a https://medium.com/@ashleygritzman/available-now-open-source... * There have been a few other efforts, all of which seem to fall short of Hinton's performance. Gritzman does a good job of covering those other efforts in his Medium article. None of these efforts propose any new ideas, as far as I can tell. RESPONSES TO #3: * Me too. So does Hinton: https://openreview.net/forum?id=HJWLfGWRb https://openreview.net/forum?id=HJWLfGWRb ... and so does everyone else. * Alas, as Paul Barham and Michal Isard at Google Brain showed earlier this year, currently it can be challenging to scale capsule networks to large datasets and output spaces, in some circumstances, due in part to current software (e.g., PyTorch, TensorFlow) and hardware (e.g., GPUs, TPUs) systems, which are highly optimized for a fairly small set of computational kernels, in a way that is tightly coupled with memory hardware, leading to poor performance on non-standard workloads, including basic operations on capsules. Source: Barham and Isard (2019) - https://dl.acm.org/citation.cfm?id=3321441 https://dl.acm.org/citation.cfm?id=3321441 (the PDF is available for free download at that link). * My draft paper mentions Barham and Isard’s work.
- p1esk 7y ago1. Both Hinton’s capsules papers have been released at the same time (Oct 2017). You can see the first comment on OpenReview page for the EM paper is dated Nov 2017. From what I remember, the two papers appear very similar with the main difference in how the routing is implemented. 2. You cite a convnet result from 2011 (!). Don’t you think a modern convnet would do vastly better on this task? 3. Could input size play a role? Did you try feeding 96x96 inputs to the models you’re comparing against, to see if they also benefit from it? 4. I’m a bit confused as to why other implementations failed to reproduce Hinton’s results given that he open sourced their code (link in the first OpenReview comment). 5. Ok, Imagenet is too slow, how about Cifar-10? What would it take to reach, say, 95%? That would be equivalent to a well trained Resnet-18. If you can show such result, I personally would become more interested, because I worked quite a bit with Cifar-10, but not with Norb. I think you might be onto something, but it’s still not clear that capsules approach is scalable and ultimately superior to plain convnets.
- fheinsen 7y agoI’m surprised you did not comment on the fact that my version of EM routing also achieves SOTA on another domain, natural language. Same code. Here are the answers to your questions: 1. The final, published version is stamped “ICLR 2018,” so I used that year. 2. I don’t know if a conventional CNN can do this with 10x fewer parameters, while also learning to do a form of “reverse graphics” without explicitly optimizing for it. (I wouldn’t know how to get a CNN to do that without explicitly making it a training objective.) 3. IIRC, the convnet model from 2011 accepts 96x96 images. As to why Hinton et al. downsample images to 9x smaller, I suspect (but don’t know for sure) they had no choice to conserve memory and computation using their version of EM routing. I was able to reduce memory and computation with my variant of EM routing (by between one and two orders of magnitude) by setting the first routing layer to accept a variable number of inputs, without regard to location in image. 4. Me too. But you asked me about work other than Hinton’s, and that’s all I could find! 5. CIFAR10 is on the to-do list (work permitting!) :-)
- p1esk 7y ago