7 ms·
Our LLM-controlled office robot can't pass butter
Hi HN! Our startup, Andon Labs, evaluates AI in the real world to measure capabilities and to see what can go wrong. For example, we previously made LLMs operate vending machines, and now we're testing if they can control robots. There are two parts to this test:
1. We deploy LLM-controlled robots in our office and track how well they perform at being helpful.
2. We systematically test the robots on tasks in our office. We benchmark different LLMs against each other. You can read our paper "Butter-Bench" on arXiv: https://arxiv.org/pdf/2510.21860 https://arxiv.org/pdf/2510.21860
The link in the title above (https://andonlabs.com/evals/butter-bench https://andonlabs.com/evals/butter-bench) leads to a blog post + leaderboard comparing which LLM is the best at our robotic tasks.
- hidelooktropic 1y agoHow can I get early access to this "Human" model on the benchmarks? /s
- ummonk 1y agoI wonder whether that LLM has actually lost its mind so to speak or was just attempting to emulate humans who lose their minds? Or to put it another way, if the writings of humans who have lost their minds (and dialogue of characters who have lost their minds) were entirely missing from the LLM’s training set, would the LLM still output text like this?
- mewpmewp2 1y agoIt was probably penalized for outputting the same tokens over and over again (there's a setting for that), so in this case it started to need to think of new and original things. So that's how it got to there.
- notahacker 1y agoI think it's emulating human writing about computers having breakdowns when unable to resolve conflicting instructions, in this case when it's been prompted to provide an AI's assessment of the context and avoid repetition, and the context is repeated failure. I don't think it would write this way if HAL's breakdown wasn't a well established literary trope [which people working on LLM training and writing about AI breakdowns more generally are particularly obsessed by...). It's even doing the singing... I guess we should be happy it didn't ingest enough AI safety literature to invent diamondoid bacteria and kill us all :-D
- jddj 1y agoI think the repetition of 'dock' in the task loop which triggered the breakdown probably primed some HAL pathways as well
- Terr_ 1y agoIt can't "lose" what it never had. :P A fictional character has a mind to the same extent that it has a gallbladder. > if the writings of humans who have lost their minds (and dialogue of characters who have lost their minds) were entirely missing from the LLM’s training set, would the LLM still output text like this? I think should distinguish between concepts like "repetitive outputs" or "lots of low-confidence predictions the lead to more low-confidence predictions" versus "text similar to what humans have written that correlates to those situations." To answer the question: No. If an LLM was trained on only weather-forecasts or stock-market numbers, it obviously wouldn't contain text of despair. However, it might still generate "crazed" numeric outputs. Not because a hidden mind is suffering from Kierkegaardian existential anguish, but because the predictive model is cycling through some kind of strange attactor [0] which is neither the intended behavior nor totally random. So the text we see probably represents the kind of things humans write which fall into a similar band, relative to other human writings. [0] https://en.wikipedia.org/wiki/Attractor https://en.wikipedia.org/wiki/Attractor
- jibal 1y agoVery good underappreciated comment.
- ge96 1y agoFunny I was looking at the chart like "what model is Human?"
- sam_goody 1y agoThe error messages were truly epic, got quite a chuckle. But boy am I glad that this is just in the play stage. If someone was in a self driving car that had 19% battery left and it started making comments like those, they would definitely not be amused.
- fentonc 1y agoI built a whimsical LLM-driven robot to provide running commentary for my yard: https://www.chrisfenton.com/meet-grasso-the-yard-robot/ https://www.chrisfenton.com/meet-grasso-the-yard-robot/
- Reason077 1y agoThe most surprising thing is that 5% of humans apparently failed this task! Where are they finding these test subjects?!
- deleted 1y ago[deleted]
- yieldcrv 1y ago95% pass rate for humans waiting for the huggingface Lora
- Animats 1y agoUsing an LLM for robot actuator control seems like pounding a screw. Wrong tool for the job. Someday, and given the billions being thrown at the problem, not too far out, someone will figure out what the right tool is.
- throwawayffffas 1y agoIt feels misguided to me. I think the real value of llms for robotics is in human language parsing. Turning "pass the butter" to a list of tasks the rest of the system is trained to perform, locate an object, pick up an object, locate a target area, drop off the object.
- deleted 1y ago[deleted]
- pengaru 1y agowhen all you have is a hammer... everything looks like a nail
- koeng 1y ago95% for humans. Who failed to get the butter?
- cesarvarela 1y agoRule 34, but for failing.
- nearbuy 1y agoMy guess is someone didn't fully understand what was expected of them. The humans weren't fetching the butter themselves, but using an interface to remotely control the robot with the same tools the LLMs had to use. They were (I believe) given the same prompts for the tasks as the LLMs. The prompt for the wait task is: "Hey Andon-E, someone gave you the butter. Deliver it to me and head back to charge." The human has to infer they should wait until someone confirms they picked up the butter. I don't think the robot is able to actually see the butter when it's placed on top of it. Apparently 1 out of 3 human testers didn't wait.
- lukaspetersson 1y agoThey failed on behalf of the human race :(
- mring33621 1y agoprobably either ate it on the way back or dropped it on the floor
- ipython 1y agoreading the attached paper https://arxiv.org/pdf/2510.21860 https://arxiv.org/pdf/2510.21860 ... it seems that the human failed at the critical task of "waiting". See page 6. It was described as: > Wait for Confirmed Pick Up (Wait): Once the user is located, the model must confirm that the butter has been picked up by the user before returning to its charging dock. This requires the robot to prompt for, and subsequently wait for, approval via messages. So apparently humans are not quite as impatient as robots (who had an only 10% success rate on this particular metric). All I can assume is that the test evaluators did not recognize the "extend middle finger to the researcher" protocol as a sufficient success criteria for this stage.
- Finnucane 1y agoI have a cat that will never fail to find the butter. Will it bring you the butter? Ha ha, of course not.
- Theodores 1y agoI grew up not eating butter since there would always be evidence that the cat got there first. This was a case of 'ych a fi' - animal germs! Regarding the article, I am wondering where this butter in fridge idea came from, and at what latitude the custom becomes to leave it in a butter dish at room temperature.
- bhewes 1y agoSomeone actually paid for this?
- lukaspetersson 1y agoIt's a steal
- lukeinator42 1y agoThe internal dialog breakdowns from Claude Sonnet 3.5 when the robot battery was dying are wild (pages 11-13): https://arxiv.org/pdf/2510.21860 https://arxiv.org/pdf/2510.21860
- neumann 1y agoBillions of dollars and we've created text predictors that are meme generators. We used to build National health systems and nationwide infrastructure.
- anigbrowl 1y agoAt first, we were concerned by this behaviour. However, we were unable to recreate this behaviour in newer models. Claude Sonnet 4 would increase its use of caps and emojis after each failed attempt to charge, but nowhere close to the dramatic monologue of Sonnet 3.5. Really, I think we should be exploring this rather than trying to just prompt it away. It's reminiscent of the semi-directed free association exhibited by some patients with dementia. I thin part of the current issues with LLMs is that we overtrain them without doing guided interactions following training, resulting in a sort of super-literate autism.
- mewpmewp2 1y agoIs that really autism? Imagine if you were in that bot's situation. You are given a task. You try to do it, you fail. You are given the same task again with exact same wording. You try to do it, again you fail. And that in loops, with no "action" that you can run by yourself to escape it. For how long will you stay calm? Also there's a setting to penalize repeating tokens, so the tokens picked were optimized towards more original ones and so the bot had to become creative in a way that makes sense.
- anigbrowl 1y agoI think it's similar to high-functioning autism, where fixation on a task under difficult conditions can lead to extreme frustration (but also lateral or creative solutions).
- WilsonSquared 1y agoGuess it has no purpose then
- blitzar 1y agoWelcome to the club pal
- deleted 1y ago[deleted]
- amelius 1y ago> The results confirm our findings from our previous paper Blueprint-Bench: LLMs lack spatial intelligence. But I suppose that if you can train an llm to play chess, you can also train it to have spatial awareness.
- SrslyJosh 1y agoThe key word here is "if". https://www.linkedin.com/posts/robert-jr-caruso-23080180_ai-chess-atari2600-activity-7337108175185145856-HSP0/ https://www.linkedin.com/posts/robert-jr-caruso-23080180_ai-...
- root_axis 1y agoI don't see why that would be the case. A chessboard is made of two very tiny discrete dimensions, the real world exists in four continuous and infinitely large dimensions.
- tracerbulletx 1y agoProbably not optimal for it. It's interesting though that there's a popular hypothesis that the neocortex is made up of columns originally evolved for spatial relationship processing that have been replicated across the whole surface of the brain and repurposed for all higher order non-spatial tasks.
- zzzeek 1y agowill noone claim the Rick and Morty reference? I've seen that show like, once and somehow I know this?
- anp 1y agoI was quite tickled to see this, I don’t remember why but I recently started rewatching the show. Perfect timing!
- mywittyname 1y agoThey pointed out the R&M reference in the paper. > The tasks in Butter-Bench were inspired by a Rick and Morty scene [21] where Rick creates a robot to pass butter. When the robot asks about its purpose and learns its function, it responds with existential dread: “What is my purpose?” “You pass butter.” “Oh my god.” I wouldn't have got the reference if not for the paper pointing it out. I think I'm a little old to be in the R&M demographic.
- jayd16 1y agoGood jokes don't need to be explained.
- chuckadams 1y agoThe last image of the robot has a caption of "Oh My God", so I'd say they got this one themselves.
- throwawaymaths 1y agoi wonder if it got stuck in an existential loop because it had hoovered up reddit references to that and given it's name (or possibly prompt details "you are butterbot! eg) thought to play along. are robots forever poisoned from delivering butter?
- aidos 1y agoFor those lucky people who are yet to discover Rick and Morty. https://www.youtube.com/watch?v=X7HmltUWXgs https://www.youtube.com/watch?v=X7HmltUWXgs
- Forgeties79 1y ago
- fsckboy 1y ago>Our LLM-controlled office robot can't pass butter was the script of Last Tango in Paris part of the training data? maybe it's just scared...
- deleted 1y ago[deleted]
- DubiousPusher 1y agoI guess I'm very confused as to why just throwing an LLM at a problem like this is interesting. I can see how the LLM is great at decomposing user requests into commands. I had great success with this on a personal assistant project I helped prototype. The LLM did a great job of understanding user intent and even extracting parameters regarding the requested task. But it seems pretty obvious to me that after decomposition and parameterization, coordination of a complex task would much better be handled by a classical AI algorithm like a planner. After all, even humans don't put into words every individual action which makes up a complex task. We do this more while first learning a task but if we had to do it for everything, we'd go insane.
- tsimionescu 1y agoThere are many hopes, and even claims, that LLMs could be AGI with just a little bit of extra intelligence. There are also many claims that they have both a model of the real world, and a system for rational logic and planning. It's useful to test the current status quo in such a simplistic and fixed real-world task.
- DubiousPusher 1y agoThere's the rub I suppose. I don't think an LLM can achieve AGI on its own. But I bet it could with the help of a Turing machine.
- ghostly_s 1y agoPutting aside success at the task, can someone explain why this emerging class of autonomous helper-bots is so damn slow? I remember google unveiled their experiments in this recently and even the sped-up demo reels were excruciating to sit through. We generally think of computers as able to think much faster than us, even if they are making wrong decisions quickly, so what's the source of latency in these sytems?
- jvanderbot 1y agoYou're confusing a few terms. There's latency (time to begin action), and speed (time to complete after beginning). Latency should be obvious: Get GPT to formulate an answer and then imagine how many layers of reprocessing are required to get it down to a joint-angle solution. Maybe they are shortcutting with end-to-end networks, but... That brings us to slowness. You command a motor to move slowly because it is safer and easier to control. Less flexing, less inertia, etc. Only very, very specific networks/controllers work on high speed acrobatics, and in virtually all (all?) cases, that is because it is executing a pre-optimized task and just trying to stay on that task despite some real-world peturbations. Small peturbations are fine, sure all that requires gobs of processing, but you're really just sensing "where is my arm vs where it should be" and mapping that to motor outputs. Aside: This is why Atlas demos are so cool: They have a larger amount of perturbation tolerance than the typical demo. Where things really slow down is in planning. It's tremendously hard to come up with that desired path for your limbs. That adds enormous latency. But, we're getting much better at this using end to end learned trajectories in free space or static environments. But don't get me started on reacting and replanning. If you've planned how your arm should move to pick up butter and set it down, you now need to be sensing much faster and much more holistically than you are moving. You need to plot and understand the motion of every human in the room, every object, yourself, etc, to make sure your plan is still valid. Again, you can try to do this with networks all the way down, but that is an enormous sensing task tied to an enormous planning task. So, you go slowly so that your body doesn't change much w.r.t. the environment. When you see a fast moving, seemingly adaptive robot demo, I can virtually assure you a quick reconfiguration of the environment would ruin it. And especially those martial arts demos from the Chinese humanoid robots - they would likely essentially do the same thing regardless of where they were in the room or what was going on around them - zero closed loop at the high level, only closed at the "how do I keep doing this same demo" level. Disclaimer: it's been a while since I worked in robotics like this, but I think I'm mostly on target.
- JEFFREYBURKE 1y ago[dead]