3 ms·
This is CLIP. Here if you would see, the model is pretrained to ingest both images and text. However if you would see the prediction mode, its basically text ge
by prats226 2y ago
This is CLIP. Here if you would see, the model is pretrained to ingest both images and text. However if you would see the prediction mode, its basically text generation only. None of the multi-modal architectures think of both image and text as first class citizens