4 ms·
"We have the face oval, two eyes, a nose and a mouth. For a CNN, a mere presence of these objects can be a very strong indicator to consider that there is a fac
by ilzmastr 9y ago
"We have the face oval, two eyes, a nose and a mouth. For a CNN, a mere presence of these objects can be a very strong indicator to consider that there is a face in the image. Orientational and relative spatial relationships between these components are not very important to a CNN."
^^^ What? The opposite of this is the mainstream I thought. The promise of DL is to learn hierarchical models of your data. The network learns edge filters, learns combinations of edge filters that differentiate an eye vs a nose, but doesn't learn combinations of intermediate features that determine a face? ppl usually say with a deep enough network an hierarchical concept can be learned...
- danharaj 9y agoThe problem is max pooling, a common technique, which destroys such information to gain some invariance in the representation.
- simonster 9y agoBeyond the initial stages of the network, current SOTA CNNs use strided convolution in addition to (Inception, NASNet) or instead of (ResNet, DenseNet) max pooling. But my impression is that this has more to do with computational efficiency than anything else. Even with max pooling, you can maintain spatial information if you construct the preceding filters properly. But what's important in the example in the post is not the absolute locations of the parts of the face, but the spatial relationships among them, and this is actually something CNNs appear to be reasonably good at handling. CNNs achieve superhuman performance in identifying faces from natural images, so I doubt that a CNN would have trouble telling apart the faces shown in the article. With that said, I believe that CNNs are merely one approach to understanding images that, given enough data, appears to work quite well. It is quite possible that, by encoding a stronger prior regarding the world into the network architecture, you can accomplish the same goals more accurately with less data. The appeal of the capsules work is that the approach is substantially different from the CNNs that have been tweaked to recognize images over the last 5 years, but still appears to achieve good (and sometimes superior) performance on difficult tasks.
- rdlecler1 9y agoIntuitively this is the idea behind using genetic algorithms encoding a generative network. This gives you a species level architecture evolved for a general class of problems which is then optimized with a learning phase for a more specific problem.
- eref 9y agoThe same problem occurs with avg pooling. Strided conv also allows to "pool" neurons in the layer below to reduce the number of neurons in subsequent layers, but, in practice, deeper neurons then also have trouble learning precise representations of the locations of the things below (but much more info is retained compared to avg/max pooling). Capsules can presumably learn such things much more accurately because they can, in principle, learn precise geometric mappings to infer positions independently of the viewpoint. However, the results so far are not much better than scalar output neurons. Capsules do perform a bit better in terms of robustness against adversarial examples and overlapping objects.
- maffydub 9y agoDo you know if anyone's looked at weighted average pooling, e.g. weighted by a Gaussian centred on the middle of the receptive field? It feels like this doesn't throw all the spatial information, but also might not be quite as hard to train as capsule networks? There are some details I haven't thought throw on this, but I'd imagine you'd want your stride length to be around the standard deviation of the Gaussian. Any pointers to papers on this (or comments on why this obviously won't work) would be very welcome - I'm still trying to develop my intuition on all this!
- eref 9y agoYou'd also lose most of the information. If there is only a single active neuron among the inputs to a Gaussian kernel neuron, you would at least have info about the distance of that to the center of the receptive field, but no directionality. If there are multiple active neurons among the inputs, you'd lose most distance-to-center info. Basically imagine avg pooling as spatial downsampling by box filter or surface area integration, and Gaussian pooling as downsampling by Gaussian filtering.
- maffydub 9y agoThanks! I agreed with the intuition around spatial downsampling. I was thinking that the next layer in the network would respond to multiple samples (i.e. convolutions of the Gaussian at different positions) and, as long as you didn't have too many active neurons on the previous layer, it could extract a measure of position. If you have too many active neurons then, as you say, you encounter aliasing effects, but I think the same is true with capsule networks - they're not expected to handle particularly high-frequency features, are they? Either way, thanks for your comment!