3 ms·MobileCLIP: Fast Image-Text Models Through Multi-Modal Reinforced Training2 points by zerojames 2y ago