4 ms·
The problem is an ANN or LLM has no physical location or “point of view”, which has led ML practitioners toward the concept of embodiment: where algorithms and
by refibrillator 2y ago
The problem is an ANN or LLM has no physical location or “point of view”, which has led ML practitioners toward the concept of embodiment: where algorithms and agents no longer learn from datasets of images, videos or text curated primarily from the internet - instead they learn through interactions with their environments from an egocentric perception similar to humans.
Curiously I didn’t see any mention of this term in your paper. But you did touch on the contrast between egocentric and allocentric representations, which lies at the heart of the issue.
I think the grand challenge of AGI will be unifying these two to reap the benefits of both while minimizing the dangers of imbalance. A bit like physics wrt gravity and quantum mechanics.
We’ll need allocentric perspective for globally optimal decisions (superorganism), but we’ll also need egocentric reasoning to enable robust autonomy in unique local environments.
Also worth mentioning that multimodal representation learning seems like it essentially achieves the sort of cross modality behavior of grid cells, although physical location (not positional encoding) is not an input in any of the research I’m aware of.
- bryan0 2y agoAre there groups actively attempting this? I was thinking that this type of data gathering is a simple solution to LLMs “running out of data” to train on. Just put a camera in a room and let the model learn by exploring its environment.