4 ms·
The model was trained on video classification, image qa and image captioning. Video captioning and video qa is not trained, yet the model shows results on those
by subho406 7y ago
The model was trained on video classification, image qa and image captioning. Video captioning and video qa is not trained, yet the model shows results on those tasks.