4 ms·
I think you are reading too far into this. The title of the technical paper is “ Video generation models as world simulators”. This is “just” a transformer tha
by grbsh 3y ago
I think you are reading too far into this. The title of the technical paper is “ Video generation models as world simulators”.
This is “just” a transformer that takes in a sequence of noisy image (video frame) tokens + prompt, and produces a sequence of less noisy video tokens. Repeat until noise gone.
The point they’re making, which is totally valid, is that in order for such a model to produce videos with realistic physics, the underlying model is forced to learn a model of physics (a “world simulation”).
- nopinsight 3y agoAlphaGo and AlphaZero were able to achieve superhuman performance due to the availability of perfect simulators for the game of Go. There is no such simulator for the real world we live in. (Although pure LLMs sorta learn a rough, abstract representation of the world as perceived by humans.) Sora is an attempt to build such a simulator using deep learning. This actually affirms my comment above. “Our results suggest that scaling video generation models is a promising path towards building general purpose simulators of the physical world.” https://openai.com/research/video-generation-models-as-world-simulators https://openai.com/research/video-generation-models-as-world... What part of my argument do you disagree about?
- lanternfish 3y ago`since it is trained to simulate the real world, as opposed to imitate the pixels.` It's not that its learning a model of the world instead of imitating pixels - the world model is just a necessary emergent phenomenon from the pixel imitation. It's still really impressive and very useful, but it's still 'pixel imitation'