4 ms·
>But how does LLMs help in making chat bots better, help with this "multi-sensory" data. All three offerings from the linked blog post are either Vision LLMs o
by nl 2mo ago
>But how does LLMs help in making chat bots better, help with this "multi-sensory" data.
All three offerings from the linked blog post are either Vision LLMs or Vision/Action LLMs.
- qsera 2mo agoAction LLMs work by generating text underneath. Just some higher level software interprets the text generated and do some action. So the immediate inference result is still text. That does not help a lot.
- ACCount37 2mo agoDepends entirely on VLA arch. Some have dedicated action diffusion heads that work in a standalone non-text action output space. Much like an LLM can either use an external TTS or have audio output heads attached to it directly for native S2S. But your entire premise is wrong regardless of that. Even if VLAs were forever bound to outputting text, you'd have to prove that they're fundamentally incapable of emitting text that maps to useful action sequences. No proof of that whatsoever - and plenty of empirical evidence suggests otherwise. Even non-specialist LLMs like ChatGPT are getting better at controlling robots and navigating 3D environments, if slowly.
- qsera 2mo ago>But your entire premise is wrong regardless of that. You don't understand what I am saying. The crux of your misunderstanding is here >emitting text that maps to useful action sequences If you have a static mapping from text to action, then you are throwing away all the advantage of using an AI. The whole point of AI is that you can get an output from an input without explicit mapping. So If you use explicit mapping anywhere in the chain, then you lose most of the advantage of using the AI. So if your hardware, physical vocabulary is limited, like move left/right/up/down then what you say could work. But something that have the dexterity of a human form, this vocabulary is nearly infinite. You won't be able to use explicit mapping there.
- ACCount37 2mo agoYou can literally have an LLM output target joint angles. As text. To be decoded by an explicit decoder, and executed by the robot. Some early VLAs did exactly that. Your entire premise is wrong. Modern action decoders are different, and usually take the form of neural networks trained end to end jointly with the rest of the model. Not fundamentally more expressive, just more in line with what we want.
- nl 2mo ago> Action LLMs work by generating text underneath This isn't true. Obviously there is a lot of variety in architecture, but in the prototypical example there are vision and languages encoders and an action decoder which decodes direction into action steps. Eg, Hugging Face SmolVLA: > Specifically, the VLM processes sensorimotor states, including images from multiple RGB cameras, and a language instruction describing the task. In turn, the VLM outputs features directly fed to the action expert, which outputs the final 3 continuous actions.[1] Or NVidia's GR00T N1: > A diffusion transformer (DiT) processes the robot’s proprioceptive state and action, which are then cross-attended with image and text tokens from the Eagle-2 VLM backbone to output the denoised motor actions.[2] (Emphasis mine) [1] https://arxiv.org/pdf/2506.01844 https://arxiv.org/pdf/2506.01844 [2] https://arxiv.org/pdf/2503.14734 https://arxiv.org/pdf/2503.14734