4 ms·
Combining embeddings is the backbone of multimodal LLMs, such as InstructBLIP[0] or LLaVA[1]. Those architectures take the output tokens from a (frozen) vision
by pkage 3y ago
Combining embeddings is the backbone of multimodal LLMs, such as InstructBLIP[0] or LLaVA[1]. Those architectures take the output tokens from a (frozen) vision transformer and train a very small projection layer between the output token space of the ViT and the input space of the LLM.
[0] https://arxiv.org/abs/2305.06500 https://arxiv.org/abs/2305.06500
[1] https://llava-vl.github.io/ https://llava-vl.github.io/
- 3abiton 3y agoI wonder when Dalle3 is released if OpenAi will release any technical documents about the integration with gpt4. Their approach might be similar.
- isaacfung 3y agoOpenAI used this approach for CLIP before blip and llava. CLIP is used to encode the text prompt in stable diffusion. Not sure about Dalle.