3 ms·
I don’t think image+text pair models (E.G. CLIP) would be very useful for AV tasks. They’re not very good at fine-grained classification or counting the number
by jowday 5y ago
I don’t think image+text pair models (E.G. CLIP) would be very useful for AV tasks. They’re not very good at fine-grained classification or counting the number of instances in a scene. Not even getting into latency or model size.