4 ms·
at the very least you could say "parsing and predicting text, images, and audio". and you would be correct - physical embodiment and spatial reasoning are missi
by stathibus 2y ago
at the very least you could say "parsing and predicting text, images, and audio". and you would be correct - physical embodiment and spatial reasoning are missing.
- ben_w 2y agoJust spatial resoning, people have already demonstrated it controlling robots.
- Yizahi 2y agoIt's all just text though, both images and audio are presented to LLM as a text, the training data is a text and all it does is append small bits of text to a larger text iteratively. So parent poster was correct.
- famouswaffles 2y ago>It's all just text though, both images and audio are presented to LLM as a text This is not true