4 ms·
Video data is still very much untapped and likely to unlock a step function worth of data. Current image-language models are trained mostly on {image, caption}
by jerpint 2y ago
Video data is still very much untapped and likely to unlock a step function worth of data. Current image-language models are trained mostly on {image, caption} pairs with a bit of extra fine tuning
- KPGv2 2y agoDo you think it matters that there's orders of magnitude less video (and audio) data than text data?
- jerpint 2y agoI’m not sure I agree, text is full of compressed information but lacks all of the visual cues we all use to navigate and understand our world. Video data also has temporal components which text is really bad at.
- mike_hearn 2y agoOpenAI has been training on YouTube for years.