4 ms·
Absolutely. There's no doubt teams already training multi-modal models on subsets of the videos in YouTube (there's already published examples training on Minec
by janekm 3y ago
Absolutely. There's no doubt teams already training multi-modal models on subsets of the videos in YouTube (there's already published examples training on Minecraft videos).
There's even tons of POV videos now literally showing human life experience in many scenarios: https://www.youtube.com/watch?v=Ipe9xJCfuTM https://www.youtube.com/watch?v=Ipe9xJCfuTM
There's enough in there to approximate the sensorial inputs that humans get through their life.