3 ms·MobileCLIP: Fast Image-Text Models Through Multi-Modal Reinforced Training1 points by zerojames 3y ago