3 ms·
My understanding is that in multimodal models, both text and image vectors align to the same semantic space, this alignment seems to be the main difference from
by netdur 2y ago
My understanding is that in multimodal models, both text and image vectors align to the same semantic space, this alignment seems to be the main difference from text-only models."