4 ms·
It is intriguing. I would have guessed human language with all its structure, would require way less parameters. If one were to look at the possibility of a 400
by random-walker 4y ago
It is intriguing. I would have guessed human language with all its structure, would require way less parameters. If one were to look at the possibility of a 400x400 image and say 1000 words that describes it, the image would be from a 160,000^(16M) dimensional space. Whereas the 1000 words would require 1000^(40000) dimensional space. Space of all possible words seems smaller than all possible images. True, that visual image has a lot of redundancies, meaning I can change a lot of pixels and still the person will say both the images are the same. Whereas if you look at language, if you change even a few characters, humans might recognize the change. But human language is very heavily structured. It is constrained by grammar, constrained by semantics ('purple banana danced on top of the super-scalar processor' is nonsense.) etc. So once you apply these, the search space seems to get much much more constrained. Images are also constrained, for example if you take a random data point in the above space, it will look like noise to us. The visually interesting sub-space is much smaller. You can even constrain by sort of stochastic visual grammar (see David Mumford's work). The idea being humans have faces, faces have eyes etc. So if you see a face, you are more likely to see co-occurring parts as well. So both of them have a more constrained space that we are really interested in (one can define your own version of this). Our training of models is to differentiate/generate them within this space. And the question is, is one of these spaces definitely much smaller than the other. I would have presumed the constrained visual space is much larger than the constrained text space. Thus my only answer to the current contradiction being, we seem to be doing better with vision models than in language models. It could partly be that since we are more sensitive to errors in output of text, thus it is harder to find simpler models.
Another way to look at this. Let's look at training data. A human child might see 65M images (assuming 1 image per sec given temporal redundancy, 10hrs awake) by age of 5. Would have heard 50-100M words (assuming 20-30k words/day) and spoken a few million words and so 'trained' for 20-40k hours. And the child can speak reasonably well by this time and detect common objects etc. Stable diffusion was trained on 170 Million images (1-2 order mag diff from child) or 3x10^13 bits of info and trained for 150K GPU hours giving a 1 Billion parameters. GPT3 was trained on 600x10^9 tokens of info and trained for 900k GPU hours giving a 170 Billion parameter model. So it seems like stable diffusion is getting a lot better compression. About 1/30k vs 1/4 compression.
Caveat: Human learning process is much more complex and more effective (as of now at-least). We also learn actively by interacting with the world by changing the world etc. Think of the child gazing at the apple and looking at it from different angles or creating gibberish sentences very close to actual sentences and getting precise adult correction. We have a model of the world and we reason about it and provide 'consistency guarantees' between various questions about it, correctness etc (again all these only to a certain extent). Try asking questions like "I have a nail on the wall that is parallel to the floor, now I hang a painting on the wall. How is the painting placed with respect to the floor". Even a child would answer this.