5 ms·
> Action LLMs work by generating text underneath This isn't true. Obviously there is a lot of variety in architecture, but in the prototypical example there a
by nl 2mo ago
> Action LLMs work by generating text underneath
This isn't true.
Obviously there is a lot of variety in architecture, but in the prototypical example there are vision and languages encoders and an action decoder which decodes direction into action steps. Eg, Hugging Face SmolVLA:
> Specifically, the VLM processes sensorimotor states, including images from multiple RGB cameras, and a language instruction describing the task. In turn, the VLM outputs features directly fed to the action expert, which outputs the final 3 continuous actions.[1]
Or NVidia's GR00T N1:
> A diffusion transformer (DiT) processes the robot’s proprioceptive state and action, which are then cross-attended with image and text tokens from the Eagle-2 VLM backbone to output the denoised motor actions.[2]
(Emphasis mine)
[1] https://arxiv.org/pdf/2506.01844 https://arxiv.org/pdf/2506.01844
[2] https://arxiv.org/pdf/2503.14734 https://arxiv.org/pdf/2503.14734