3 ms·
> I have this diffuse idea in my head that most possible images do not occur in the real world and that there are way more degrees of freedom into direction tha
by pakl 10y ago
> I have this diffuse idea in my head that most possible images do not occur in the real world and that there are way more degrees of freedom into direction that just don't occur in the real world but this idea is just too diffuse so that I am currently unable to pin and write down.
Yes! You're on the right track! The number of degrees of freedom of images of pixels and textures is HUGE. There is not enough data to practically learn directly from those images. So the deep networks are starved for data -- even with the big datasets they are trained on. (It's only thanks to the way they are set up they do well when tested on very similar images, like sharp hi-res photos. But they fail to generalize to other kinds of images.)
So how can you reasonable reduce these degrees of freedom?
It turns out that the continuity of reality itself provides a powerful constraint that can reduce the degrees of freedom. See, when a ball rolls along, this physical event is not just a collection of textures to be memorized. It's an ordered sequence of textures that vary in a consistent and regular way because of many learnable physical constraints (like lighting).
So, it turns out you can reduce the dimensionality by making a particular kind of large recurrent neural net learn to predict the future in video. Our very preliminary testing shows it works shockingly well.