3 ms·
I agree. We've [potentially] given the grounding assimilation step a boost already, since language is already organized. Imagine that a language model is full
by Syntonicles 4y ago
I agree. We've [potentially] given the grounding assimilation step a boost already, since language is already organized.
Imagine that a language model is fully integrated with sense-data that exceeds human first-hand experience. Perhaps they are trained on and can generate realistic 3D models of objects, and derive estimates of their internal construction, weight, etc. Perhaps they recall infrared emissions or opacity to EM wavelengths. Would we truly "know" what we're talking about by that standard?
I'm not actually sure why we don't consider generative image models to be grounded already. They seem to be able to modify, transform and rotate imagery. That indicates spatial understanding to me, and I'm not sure how much more we must require of them without having to exclude blind or otherwise disabled humans from our definition of comprehension.