9 ms·
I feel LeCun got roped in debating the likes of Marcus and Yudkowsky. This has made his arguments lose nuance and become rigid. I also can't escape the feeling
by bonzaidrinkingb 3y ago
I feel LeCun got roped in debating the likes of Marcus and Yudkowsky. This has made his arguments lose nuance and become rigid. I also can't escape the feeling that if Facebook was tuned into Transformers, they would have shipped earlier, so there must have been some resistance or underestimation that's now repeated "They can't reason", "They can't plan", "They can't understand the world", "They are a distraction / side road to AGI".
It is kind of ironic that researchers who claim LLMs lack adaptive intelligence seemingly refuse to adapt their intelligence to LLMs. If even GPT-3 can find logical holes or oversimplification in your arguments about GPTs, at one point this starts becoming embarrassing and unbecoming.
> The generation of mostly realistic-looking videos from prompts does not indicate that a system understands the physical world.
While arguably true, it also does not indicate that a system does not understand the physical world (reflections, collision detection, gravity, object permanence, long-term scene coherence, etc.).
If LeCun wants to argue it does not understand the physical world, he should do so directly. Not attack something that is not directly stated, but rather convincingly and tentatively demo'd (I myself find it hard to argue that a system that generates novel pond reflections has not memorized/stored in weights some generalization program to apply to realistic scene generation).
This demo shows it is not even a wild prediction to guess that soon (consumer tech) we will be able to discuss visual scenes with conversational AIs.
- PoignardAzur 3y ago> long-term scene coherence, FWIW none of the video models released so far demonstrate any object coherence whatsoever, which suggests they don't have the higher level capabilities you mention yet. In Sora, as soon as an object is obstructed by an obstacle or goes offscreen, it's likely to disappear or be radically transformed.
- bonzaidrinkingb 3y agoYou've seen the demos of a couple holding hands and walking, or the museum shots where all the paintings maintain coherence, or a woman temporarily obscuring a street sign. Or you haven't seen those demos. Either way...
- kgwgk 3y agoIf the "couple holding hands and walking" one is the "Beautiful, snowy Tokyo city is bustling. ..." look at the traffic on the left side of the frame: https://www.youtube.com/watch?v=ezaMd4l_5kw https://www.youtube.com/watch?v=ezaMd4l_5kw We also have the spontaneous creation and annihilation of wolves and the shape-shifting chair: https://www.youtube.com/watch?v=jspYKxFY7Sc https://www.youtube.com/watch?v=jspYKxFY7Sc https://www.youtube.com/watch?v=lfbImB0_rKY https://www.youtube.com/watch?v=lfbImB0_rKY
- hackerlight 3y agoIf we are talking analogies, this is just Sora forgetting because of limitations of how the network handles the autoregressive dynamics. When they make a bigger version of Sora this will happen less. Sora aleady has unprecedented object permanence, see the woman walking in Tokyo scene where signs and people are reconstructed after two seconds of occlusion. Soon we will have object permanence following ten or more seconds of occlusion. Then a minute. Then three minutes. Then we will figure out a trick to store long term memory. What will people say then?
- orwin 3y agoOur brain can't work that in our long-term memory btw, that why each time we remember something, we change minor aspects of said thing.
- kgwgk 3y agoIt's also what happens when we dream: everything is fluid. Things appear and disappear, people and places become someone or somewhere else, reading is difficult and hands are distorted.
- namaria 3y agoBecause these systems are dreaming about their datasets. Or hallucinating about it, as people have decided to call lately. I won't say this is a dead end. I will say we are very, very short of any sort of actual intelligence.