2 ms·For now we're using CLIP. We've also done testing with Siglip and Gemma for a full-blown vision model.by correa_brian 1y agoFor now we're using CLIP. We've also done testing with Siglip and Gemma for a full-blown vision model.