3 ms·
I think it's the fact that you can subtly pick up the way Dall-e has begun to understand the similarity between different abstract forms. Dall-e has likely intu
by planetsprite 4y ago
I think it's the fact that you can subtly pick up the way Dall-e has begun to understand the similarity between different abstract forms. Dall-e has likely intuited that there is a general similarity between the faces of different animals, or a general similarity between different organic objects and different artificial ones. We are seeing the expression of those connections the model has made: compression in abstract representation that triggers the same uneasy feeling human brains get when on a bad trip and they see faces in clouds and shadows.
The training of these image generation models minimizes the amount of necessary new visual information it needs to remember with each new image it sees during training, so whenever it sees a new image-caption pair it abstracts its qualities into an amalgamation of what its already seen. This is how Stable Diffusion managed to fit the complete corpus of human visual memory into 4.2 gigabytes.
Everything is more closely connection than the picture would imply at first glace, like how in Super Mario Bros. the bushes are just clouds colored green. Compression and efficiency leads to a hyperconnectivity in the way DALL-E ties prompts and images. This is why those old deep-dream images where everything looked like a dog's face were so unsettling. This triggers an uncanny response in the human brain, because it sees the vague shape of a face in a tree, or the vague shape of a body in the structure of a building.