4 ms·
You can literally have an LLM output target joint angles. As text. To be decoded by an explicit decoder, and executed by the robot. Some early VLAs did exactly
by ACCount37 2mo ago
You can literally have an LLM output target joint angles. As text. To be decoded by an explicit decoder, and executed by the robot. Some early VLAs did exactly that.
Your entire premise is wrong.
Modern action decoders are different, and usually take the form of neural networks trained end to end jointly with the rest of the model. Not fundamentally more expressive, just more in line with what we want.
- qsera 2mo agoYou are repeating "It can be done, trust me!". But that is not very convincing.
- ACCount37 2mo agoI'm repeating "what you claim to be impossible was done 3 years ago and was already replaced with better versions of the same idea and you are hilariously out of touch".
- qsera 2mo agoShow me one video where a robot follows textual prompts and come up with its own movements to solve the prompt.
- ACCount37 2mo agoThat's just about any video of any VLA ever. Including Gemini Robotics 2.
- qsera 2mo agoShould be trivial to link to one then...
- ACCount37 2mo agoOff the top of my head: https://www.pi.website/blog/pi07 https://www.pi.website/blog/pi07
- qsera 2mo agoAs I suspected you are fooled by this video and imagine it to be capable of much more than what is shown. This video is pretty non-marketing and is quite straight to the point. But that does not prevent you from being awed! So What is LLM is used here for? It is used for mere translation between different robots. So it is mostly symbolic translation. What I am talking about is to translation LLM inference directly to movements. For example, if you ask an LLM, how do I open the microwave door? It will list the steps. I am talking about a system that can go from "put the thing in the microwave", to action steps, without having to never once demonstrate it physically, and do it just from LLM inference. In short, the way LLMs used here is not (categorically) the way I was asking about.
- ACCount37 2mo agoRead. The. Papers. https://arxiv.org/pdf/2505.23705 https://arxiv.org/pdf/2505.23705 https://www.pi.website/download/pistar06.pdf https://www.pi.website/download/pistar06.pdf https://www.pi.website/download/pi07.pdf https://www.pi.website/download/pi07.pdf The thing literally has a diffusion "action expert" sit in the same attention system as a pre-trained VLM. And the VLM itself is ALSO trained to generate raw actions as a part of the training recipe (the first paper) - it just doesn't do it at inference time. What the "action expert" does is parallelize the action generation process - based on VLM's internal states. It's exactly the thing you claimed to be impossible. Described in detail in a paper from 2025. What's your excuse?
- qsera 2mo agoI have not overlooked anything. I had imagined that this "mapping", to have any power, would also need to be handled by an LLM (action expert). But here is the problem with that. That would not be as "intelligent" as an LLM....And you can't make it as smart as the LLM because there is not a similarly huge training data on which LLMs are trained on..
- nl 2mo agohttps://huggingface.co/blog/lerobot-release-v060#molmoact2 https://huggingface.co/blog/lerobot-release-v060#molmoact2 Read the command line prompt: --task="pick up the red cube" There is a gif directly below it. This is a completely open source model and arm you can replicate yourself.