3 ms·
I do not understand these comments at all. Sora was trained on billions of frames from video and images - they were tagged with words like "ballistic missile la
by beardedwizard 3y ago
I do not understand these comments at all. Sora was trained on billions of frames from video and images - they were tagged with words like "ballistic missile launch" and "cinematic shot" and it simply predicts the pixels like every other model. It stores what we showed it, and reproduces it when we ask - this has nothing to do with understanding and everything to do with parroting. The fact that it's now a stream of images instead of just 1 changes nothing about it.
- ninetyninenine 3y agoWhat is the difference between a machine that for all intents and purposes appears to understand something to a degree of 100 percent versus a human? Both the machine and the human are a black box. The human brain is not completely understood and the LLM is only trivially understood at a high level through the lens of stochastic curve fitting. When something produces output that imitates the output related to a human that we claim "understands" things that is objectively understanding because we cannot penetrate the black box of human intelligence or machine intelligence to determine further. In fact in terms of image generation the LLM is superior. It will generate video output superior to what a human can generate. Now mind you the human brain has a classifier and can identify flaws but try watching a human with Photoshop to try to even draw one frame of those videos.. it will be horrible. Does this indicate that humans lack understanding? Again, hard to answer because we are dealing with black boxes so it's hard to pinpoint what understanding something even means. We can however set a bar. A metric. And we can define that bar as humans. all humans understand things. Any machine that approaches human input and output capabilities is approaching human understanding.
- Jensson 3y ago> What is the difference between a machine that for all intents and purposes appears to understand something to a degree of 100 percent versus a human? There is no such difference, we evaluate that based on their output. We see these massive model make silly errors that nobody who understands it would make, thus we say the model doesn't understand. We do that for humans as well. For example, for Sora in the video with the dog in the windos, we see the dog walk straight through the window shutters, so Sora doesn't understand physics or depth. We also see it drawing the dogs shadow on the wall very thin, much smaller than the dog itself, it obviously drew that shadow as if it was cast on the ground and not a wall, it would have been very large shadow on that wall. The shadows from the shutters were normal, because Sora are used to those shadows being on a wall. Hence we can say Sora doesn't understand physics or shadows, but it has very impressive heuristics about those, the dog accurately places its paws on the platforms etc, and the paws shadows were right. But we know those were just basic heuristics since the dog walked through the shutters and its body cast shadow in the wrong way meaning Sora only handles very common cases and fails as soon as things are in an unexpected envionment.
- ninetyninenine 3y ago>There is no such difference, we evaluate that based on their output. We see these massive model make silly errors that nobody who understands it would make, thus we say the model doesn't understand. We do that for humans as well. Two things. We also see the model make things that are correct. In fact the mistakes are a minority in comparison to what it got correct. That is in itself an indicator of understanding to a degree. The other thing is, if a human tried to reproduce that output according to the same prompt, the human would likely not generate something photorealistic and the thing a human comes up with will be flawed, ugly disproportionate wrong and an artistic travesty. Does this mean a human doesn't understand reality? No. Because the human generates worse output visually than an LLM we cannot say the human doesn't understand reality. Additionally the majority of the generated media is correct. Therefore it can be said that the LLM understands the majority of the task it was instructed to achieve. Sora understands the shape of the dog. That is in itself remarkable. I'm sure with enough data sora can understand the world completely and to a far greater degree than us. I would say it's uncharitable to say sora doesn't understand physics when it gets physics wrong, and that for the things it gets right it's only heuristics.
- beardedwizard 3y agoHow can it possibly understand physics when the training data does not teach it or contain the laws of physics?
- ninetyninenine 3y agoVideo data contains physics. Objects in motion obey the laws of physics. Sora understand physics the same way you understand it.
- beardedwizard 3y agoI understand physics because science has performed a series of measurable experiments over 100s of years resulting in concrete mathematical formulas and theories that explain the laws of physics so that they can be reproduced by machines. Sora has zero of this knowledge. This is very much allegory of the cave[1]. If Sora sees a series of images that contain impossible physics, for example MC Escher paintings, what will happen? 1: https://en.wikipedia.org/wiki/Allegory_of_the_cave https://en.wikipedia.org/wiki/Allegory_of_the_cave