3 ms·
Depends entirely on VLA arch. Some have dedicated action diffusion heads that work in a standalone non-text action output space. Much like an LLM can either use
by ACCount37 2mo ago
Depends entirely on VLA arch. Some have dedicated action diffusion heads that work in a standalone non-text action output space. Much like an LLM can either use an external TTS or have audio output heads attached to it directly for native S2S.
But your entire premise is wrong regardless of that.
Even if VLAs were forever bound to outputting text, you'd have to prove that they're fundamentally incapable of emitting text that maps to useful action sequences. No proof of that whatsoever - and plenty of empirical evidence suggests otherwise. Even non-specialist LLMs like ChatGPT are getting better at controlling robots and navigating 3D environments, if slowly.
- qsera 2mo ago>But your entire premise is wrong regardless of that. You don't understand what I am saying. The crux of your misunderstanding is here >emitting text that maps to useful action sequences If you have a static mapping from text to action, then you are throwing away all the advantage of using an AI. The whole point of AI is that you can get an output from an input without explicit mapping. So If you use explicit mapping anywhere in the chain, then you lose most of the advantage of using the AI. So if your hardware, physical vocabulary is limited, like move left/right/up/down then what you say could work. But something that have the dexterity of a human form, this vocabulary is nearly infinite. You won't be able to use explicit mapping there.
- ACCount37 2mo agoYou can literally have an LLM output target joint angles. As text. To be decoded by an explicit decoder, and executed by the robot. Some early VLAs did exactly that. Your entire premise is wrong. Modern action decoders are different, and usually take the form of neural networks trained end to end jointly with the rest of the model. Not fundamentally more expressive, just more in line with what we want.
- qsera 2mo agoYou are repeating "It can be done, trust me!". But that is not very convincing.
- ACCount37 2mo agoI'm repeating "what you claim to be impossible was done 3 years ago and was already replaced with better versions of the same idea and you are hilariously out of touch".
- qsera 2mo agoShow me one video where a robot follows textual prompts and come up with its own movements to solve the prompt.
- ACCount37 2mo agoThat's just about any video of any VLA ever. Including Gemini Robotics 2.
- qsera 2mo agoShould be trivial to link to one then...
- ACCount37 2mo agoOff the top of my head: https://www.pi.website/blog/pi07 https://www.pi.website/blog/pi07
- qsera 2mo agoAs I suspected you are fooled by this video and imagine it to be capable of much more than what is shown. This video is pretty non-marketing and is quite straight to the point. But that does not prevent you from being awed! So What is LLM is used here for? It is used for mere translation between different robots. So it is mostly symbolic translation. What I am talking about is to translation LLM inference directly to movements. For example, if you ask an LLM, how do I open the microwave door? It will list the steps. I am talking about a system that can go from "put the thing in the microwave", to action steps, without having to never once demonstrate it physically, and do it just from LLM inference. In short, the way LLMs used here is not (categorically) the way I was asking about.