4 ms·
They probably use something similar to Kosmos-1 (https://arxiv.org/abs/2302.14045 https://arxiv.org/abs/2302.14045): Encode images as vectors with something lik
by benob 4y ago
They probably use something similar to Kosmos-1 (https://arxiv.org/abs/2302.14045 https://arxiv.org/abs/2302.14045): Encode images as vectors with something like CLIP, then map them to the token space and input them between <image> </image> tags.
- frabcus 4y agoSo presumably the model could output tokens that represent images as well? For multi-modal training data, e.g. HTML pages or PDFs, does the training data interleave the image tokens amongst the text tokens in the same document? Slightly limited, as doesn't get juxtaposition to text in complex ways, just linear placement of images.
- jjoonathan 4y agoIt looks like the ViT embedding is a projection, so it's not trivially reversible. I bet it could be used to guide a diffusion model or something though.