5 ms·
It's not really surprising given what we now know about autoregressive modeling with transformers. It's essentially a game of predict hidden information given v
by thunderbird120 6y ago
It's not really surprising given what we now know about autoregressive modeling with transformers. It's essentially a game of predict hidden information given visible information. As long as the relationship between the visible and hidden information is non-random you can train the model to understand an amazing amount about the world by literally just predicting the next token in a sequence given all the previous ones.
I'm curious if they do a backward pass here, would probably have value. They seem to describe sticking the text tokens first meaning that once you start generating image tokens all the text tokens are visible. That would have the model learning to generate an image with respect to a prompt but you could also literally just reverse the order of the sequence to have the model also learn to generate prompts with respect to the image. It's not clear if this is happening.
- minimaxir 6y agoThat approach wouldn't work out of the box; it sees text for the first 256 tokens and images for the following 1024 tokens, and tries to predict the same. It likely would not have much to go on if you gave it the 1024 tokens for the image and then 256 for the text later since it doesn't have much of a basis. A network optimizing for both use cases (e.g. the training set is half 256 + 1024, half 1024 + 256) would likely be worse than a model optimizing for one of the use cases, but then again models like T5 argue against it.
- lukeplato 6y agoIs this kind of happening with the CLIP classifier [1] to rank the generated images? > Similar to the rejection sampling used in VQVAE-2, we use CLIP to rerank the top 32 of 512 samples for each caption in all of the interactive visuals. This procedure can also be seen as a kind of language-guided search16, and can have a dramatic impact on sample quality. > CLIP pre-trains an image encoder and a text encoder to predict which images were paired with which texts in our dataset. We then use this behavior to turn CLIP into a zero-shot classifier. We convert all of a dataset’s classes into captions such as “a photo of a dog” and predict the class of the caption CLIP estimates best pairs with a given image. [1] https://openai.com/blog/clip/ https://openai.com/blog/clip/