4 ms·
Images are tokenized. Rumor and greatest likelihood is that it's a ViT.
by alpineidyll3 4y ago
Images are tokenized. Rumor and greatest likelihood is that it's a ViT.
- benob 4y agoThey probably use something similar to Kosmos-1 (https://arxiv.org/abs/2302.14045 https://arxiv.org/abs/2302.14045): Encode images as vectors with something like CLIP, then map them to the token space and input them between <image> </image> tags.
- frabcus 4y agoSo presumably the model could output tokens that represent images as well? For multi-modal training data, e.g. HTML pages or PDFs, does the training data interleave the image tokens amongst the text tokens in the same document? Slightly limited, as doesn't get juxtaposition to text in complex ways, just linear placement of images.
- jjoonathan 4y agoIt looks like the ViT embedding is a projection, so it's not trivially reversible. I bet it could be used to guide a diffusion model or something though.
- loufe 4y agoFor anyone else curious what ViT is: >https://en.wikipedia.org/wiki/Vision_transformer https://en.wikipedia.org/wiki/Vision_transformer >https://huggingface.co/docs/transformers/model_doc/vit https://huggingface.co/docs/transformers/model_doc/vit
- flangola7 4y agoWhat do image tokens look like? Groups of pixels?
- jjoonathan 4y agoThe ViT paper has the details: project 16x16 groups of pixels through a learned embedding, combine with positional encoding, and feed to an attention layer. It's delightful that this is practically identical to the NLP architecture with only the tiniest adaptive tweak!