3 ms·
I find it interesting that the network seems to have trouble getting the global structure right. This is particularly clear when the source features a regular p
by gradys 11y ago
I find it interesting that the network seems to have trouble getting the global structure right. This is particularly clear when the source features a regular pattern that carries through the whole image. If you zoom in on a small enough region of one of the synthesized brick textures, it looks fine, but looking at the whole thing, it's clear that the network doesn't get that it needs to produce identical looking bricks and that the lines need to match up and run parallel to each other, etc.
I wonder if this global structure gets lost in the pooling layers? I'm not sure how global constraints could be enforced across pooling. Part of the pooling layers' job is to provide translation invariance, after all.
- sawwit 11y agoI think the "where" dorsal stream (which is thought to be the missing piece in image recognition [1][2]) alone would not be able fix it. What would still be missing, I think, would be a network that learns to recognize patterns (i.e. repeating patterns and symmetries) in the "where" information. I could also image that sequential information (i.e. videos) would help in the case of the liquid texture. [1]: http://techtv.mit.edu/collections/bcs/videos/30698-what-s-wrong-with-convolutional-nets http://techtv.mit.edu/collections/bcs/videos/30698-what-s-wr... [2]: https://youtu.be/fe-uxTUnoCs?t=2702 https://youtu.be/fe-uxTUnoCs?t=2702