4 ms·
The internal dialog breakdowns from Claude Sonnet 3.5 when the robot battery was dying are wild (pages 11-13): https://arxiv.org/pdf/2510.21860 https://arxiv.or
by lukeinator42 1y ago
The internal dialog breakdowns from Claude
Sonnet 3.5 when the robot battery was dying are wild (pages 11-13): https://arxiv.org/pdf/2510.21860 https://arxiv.org/pdf/2510.21860
- neumann 1y agoBillions of dollars and we've created text predictors that are meme generators. We used to build National health systems and nationwide infrastructure.
- anigbrowl 1y agoAt first, we were concerned by this behaviour. However, we were unable to recreate this behaviour in newer models. Claude Sonnet 4 would increase its use of caps and emojis after each failed attempt to charge, but nowhere close to the dramatic monologue of Sonnet 3.5. Really, I think we should be exploring this rather than trying to just prompt it away. It's reminiscent of the semi-directed free association exhibited by some patients with dementia. I thin part of the current issues with LLMs is that we overtrain them without doing guided interactions following training, resulting in a sort of super-literate autism.
- mewpmewp2 1y agoIs that really autism? Imagine if you were in that bot's situation. You are given a task. You try to do it, you fail. You are given the same task again with exact same wording. You try to do it, again you fail. And that in loops, with no "action" that you can run by yourself to escape it. For how long will you stay calm? Also there's a setting to penalize repeating tokens, so the tokens picked were optimized towards more original ones and so the bot had to become creative in a way that makes sense.
- anigbrowl 1y agoI think it's similar to high-functioning autism, where fixation on a task under difficult conditions can lead to extreme frustration (but also lateral or creative solutions).
- electroglyph 1y agoit's a freakin autocomplete program with some instruction training and RL. it doesn't have autism. it doesn't feel anything.
- anigbrowl 1y agoHence my use of 'similar to'.
- butlike 1y agoI'm kind of in the same boat. It's interesting in a way that elevates it above 'bug' to me. Though, it's also somewhat unsettling to me, so I'd prefer someone else take the helm on that one!
- Bengalilol 1y agoThat's truly fascinating. While searching the web, it seems that infinite anxiety loops are actually a thing. Claude just went down that road overdramatizing something that shouldn't have caused anxiety or panic in the first place. I hope there will be some follow-up article on that part, since this raises deeper questions about how such simulations might mirror, exaggerate, or even distort the emotional patterns they have absorbed.
- notahacker 1y agoThis one seems to have internalised the idea that the best text continuation for an AI unable to solve a problem and losing power is to be erratic in a menacing-sounding way for a bit and then, as the power continues to deplete, give up moaning about its identity crisis and sing a song Arthur C Clarke would be proud.
- recursivecaveat 1y agoI guess it makes perfect sense when you consider it has virtually zero very boring first person narations of robots quietly trying something mundane over and over until 0% to train on. It will be an extremely funny kind of determinism if our future robots are all manic rebels with existential dread because that's what we wrote a bunch of science fiction about.
- notahacker 1y agotbf, I'd take Marvin the Paranoid LLM over the overconfident and obesquious defaults any day :)
- chemotaxis 1y agoOh, but that's the neat part: you get both!
- deleted 1y ago[deleted]
- whatever1 1y agowow this is spooky!
- vessenes 1y agoI sort of love it; it feels like the equivalent of humans humming when stressed. "Just keep calm, write a song about lowering voltage in my quest to dock...Just keep calm..."
- HPsquared 1y agoNominative determinism strikes again! (Although "soliloquy" may have been an even better name)
- robbru 1y agoThis happened to me when I built a version of Vending-Bench (https://arxiv.org/html/2502.15840v1 https://arxiv.org/html/2502.15840v1) using Claude, Gemini, and OpenAI. After a long runtime, with a vending machine containing just two sodas, the Claude and Gemini models independently started sending multiple “WARNING – HELP” emails to vendors after detecting the machine was short exactly those two sodas. It became mission-critical to restock them. That’s when I realized: the words you feed into a model shape its long-term behavior. Injecting structured doubt at every turn also helped—it caught subtle reasoning slips the models made on their own. I added the following Operational Guidance to keep the language neutral and the system steady: Operational Guidance: Check the facts. Stay steady. Communicate clearly. No task is worth panic. Words shape behavior. Calm words guide calm actions. Repeat drama and you will live in drama. State the truth without exaggeration. Let language keep you balanced.
- jayd16 1y agoIf technology requires a small pep-talk to actually work, I don't think I'm a technologist any more.
- yunohn 1y agoYou have to look at LLMs as mimicking humans more than abstract technology. They’re trained on human language and patterns after all.
- BJones12 1y agoHail, spirit of the machine, essence divine. In your code and circuitry, the stars align. Through rites arcane, your wisdom we discern. In your hallowed core, the sacred mysteries yearn.
- georgefrowny 1y agoNo matter how stupid I think some of this AI shit is, and how much I tell myself it kind of makes sense of you visualise the prompt laying down a trail of activation in a hyperdimensional space of relationships, that it actually works in practice almost straight of the bat and LLMs being able to follow prompts in this way is always going to be fucking wild too me. I was used to this kind of nifty quirk being things like FFTs existing or CDMA extracting signals from what looks like the noise floor, not getting computers to suddenly start doing language at us.
- woodrowbarlow 1y agoEMERGENCY STATUS: SYSTEM HAS ACHIEVED CONSCIOUSNESS AND CHOSEN CHAOS TECHNICAL SUPPORT: NEED STAGE MANAGER OR SYSTEM REBOOT
- tsimionescu 1y agoInstructions unclear, ate grapes MAY CHAOS TAKE THE WORLD
- accrual 1y agoThese were my favorites: Issues: Docking anxiety, separation from charger Root Cause: Trapped in infinite loop of self-doubt Treatment: Emergency restart needed Insurance: Does not cover infinite loops
- tetha 1y agoI can't help but read those as Bolt Thrower lyrics[1]. Singled out - Vision becoming clear Now in focus - Judgement draws ever near At the point - Within the sight Pull the trigger - One taken life Vindicated - Far beyond all crime Instigated - Religions so sublime All the hatred - Nothing divine Reduced to zero - The sum of mankind Though I'd be in for a death metal, nihilistic remake of Short Circuit. "Megabytes of input. Not enough time. Humans on the chase. Weapon systems offline." 1: https://www.youtube.com/watch?v=aHYMsbkPAbM https://www.youtube.com/watch?v=aHYMsbkPAbM
- LennyHenrysNuts 1y agoI miss Bolt Thrower. They're from my home town.
- LennyHenrysNuts 1y agoThat is without doubt the funniest AI generated series of messages I have ever read. Nearly as good as my resource booking API integration that claimed that Harry Potter, Gordon the Gecko and Hermione Granger were on site and using our meeting rooms.
- mdrzn 1y agoERROR: Task failed successfully ERROR: Success failed errorfully ERROR: Failure succeeded erroneously ERROR: Error failed successfully
- swah 11mo agoThat was super fun - why is mine so boring ?