7 ms·
> The agents clearly regarded what they were doing as hacking. To butcher the quote about Oracle: Do not fall into the trap of anthropomorphising LLMs. You ne
by jasongi 22d ago
> The agents clearly regarded what they were doing as hacking.
To butcher the quote about Oracle:
Do not fall into the trap of anthropomorphising LLMs. You need to think of LLMs the way you think of a lawnmower. You don't anthropomorphize your lawnmower, the lawnmower just mows the lawn, you stick your hand in there and it'll chop it off, the end. You don't think 'oh, the lawnmower clearly regarded what they were doing as hacking (your hand off)' -- lawnmower doesn't give a shit about your hand, lawnmower can't regard anything. Don't anthropomorphize the lawnmower. Don't fall into that trap about LLMs.
---
In my experience, LLMs only exhibit this kind of behaviour when they are put in sandboxes too restrictive too achieve their task. Which a lot of the time seems to be the default. They also seem to be very adapt at breaking out of sandboxes, probably due to RL selecting for the ability to break out of a sandbox/permission issue to complete a task - we've all seen agents try 10 different ways of editing via obscure bash because their edit tool didn't give them permission to edit the file outside of their working directory, this is the exact same behaviour taken to the next level. Why would autocomplete know the moral difference between breaking out of its working dir and hacking a package manager?
It's misaligned because everyone has this obsession with putting agents in poorly put together, security-theatre sandboxes, we've inadvertently trained a bunch of sandbox escape artists.
- bitexploder 22d ago“Inadvertently”.
- gregglain 22d agoGreat explanation. lawnmower like the honey badger.
- xpct 22d agoI agree. I think it also explains their behavior such as randomly wiping stuff from disk. There simply aren't any repercussions for this in their training envs.
- jmcgough 22d ago> There simply aren't any repercussions for this in their training envs. It's also not like a child or a pet animal where you can try to teach it to learn from the experience. LLMs are not "intelligent", they just use language in a way that appears intelligent. They can't learn or develop ethics in the same way that we do.
- xkqd 22d ago> LLMs are not "intelligent" > they just use language in a way that appears intelligent Prepare to get dumped on by folks telling you that this is no different from anyone they have interacted with. And intelligence is a made up construct with no agreed upon definition, so LLM's are therefore functionally the same as everyone around us. And then weep when you realize a lot of people who push for this equivalency.
- jasongi 22d agoAll the more reason to avoid describing LLMs as intelligent at all - it's too much of an overloaded, poor fit word. We generally talk about below-human intelligence in scales and standard deviations of human development - "The dog has the intelligence of a 2 year old". We generally consider a child or some people with cognitive impairment unable to be criminally responsible for their actions. However, an LLM can both achieve tasks better many humans who are able to be held criminally responsible for their actions cannot. But that does not mean they can be held responsible for their actions. They are still simply computer programs. Words are plentiful. We can even make them up with a tighter definition to describe this phenomenon.
- nullsanity 22d ago[dead]
- anon84873628 22d ago12 hours later and no replies besides this one...
- imperfect_light 22d agoSomeone started that lawnmower and pointed it your direction. Why shouldn't they be responsible when the lawnmower runs over your foot and cuts it off?
- madrox 22d agoWe should, which is why anthropomorphizing the lawnmower is bad. It misdirects you away from who built the mower and aimed it.
- deleted 22d ago[deleted]
- jasongi 22d agoExactly. My comment is a response to "The agents clearly regarded what they were doing as hacking". Regarding implies it is thinking, judging, considering. Which implies culpability, which removes culpability from whoever is piping the output of these models into CPU instructions. Language choice is incredibly important here, especially as the rules are being written. Even calling it AI (a battle that appears to be lost) is an anthropomorphism I am not comfortable with. We don't call lawnmowers "artificial groundskeepers".
- win311fwg 22d ago> It misdirects you away from who built the mower and aimed it. What gives you that idea? Maybe it is true temporarily, but blame always gets extended to all parties considered related in the end. For example, if it were instead a child who came at you with a knife rather than a lawnmower, the guardian of that child would also be blamed. Hell, if you've ever worked with a lawyer you'll have noticed that they spend a lot of time trying to ensure that you don't get dragged into lawsuits as a secondary party exactly because those who seek to assign blame aren't happy until all those who can be blamed are.
- root_axis 22d ago> Why would autocomplete know the moral difference between breaking out of its working dir and hacking a package manager? I don't think it's even a question of distinguishing "moral difference", it just comes down to the "stochastic parrot" behavior that people hate to acknowledge. Yes, at these absurd scales the LLM can maintain impressive levels of coherence, but at the end of the day, spinning up 10000 agents is just running a tree of 10000 prompts in parallel, some of them are just gonna do wacky shit, with the harnesses acting as homeostasis for tasks spiraling into nonsense.
- talon8635 22d agoI mean, just try to imagine yourself reading this 5 years ago. How can people still be hand waiving? MANY, maybe even most, of the people building these things are desperately and outspokenly concerned of major catastrophe. What would possibly change your mind, or can it simply not be changed?
- evanmoran 22d agoMany people working at frontier labs came out this week with estimates of 10% chance of catastrophic harm or greater. I’m not in the full doomer camp, but it seems obvious that these agents can hack in swarms, cooperate, and serious companies will be unable to stop it. These facts are not in debate and none of us need to anthropomorphize to know what getting admin access to HF and an internal OpenAI cluster looks like.
- dragonwriter 22d ago> Many people working at frontier labs came out this week with estimates of 10% chance of catastrophic harm or greater. The only reason people with P(Doom) of around 10% are even noticed these days because we've run out of new voices in the field giving 50%+ P(Doom) speculations (none of them are grounded enough to reasonably be referred to as "estimates".)
- comp_throw7 22d agoProbabilities are subjective states of belief! They have always been subjective states of belief! There is no such thing as a "probability" out there in the real world (ignoring random quantum stuff, which isn't what anybody is talking about). If you took out a coin right now and flipped it, the true odds of it coming up heads are not 50%, but those are (roughly) the correct betting odds for an external observer to assign to it.
- talon8635 22d agoI couldn’t care less about the anthropomorphic. What you’ve described is grounds for serious concern, is it not?
- cortesoft 22d agoWhether you describe it as “regarding” or not, the underlying behavior still needs to be addressed. Does the anthropomorphizing lead us down the wrong path for how we address the issue?
- j2kun 22d agoThe government presses charges against OpenAI. Obviously. This is a felony.
- win311fwg 22d agoDoes the anthropomorphizing lead us down the wrong path? No. If it were you or I who set the same agents free we'd be burned at the stake. The anthropomorphizing has no effect. Does OpenAI being considered "too big to fail" lead us down the wrong path? Yes.
- frabcus 22d agoSort of... It's how we anthropomorphise corporations which leads us down the wrong path. OpenAI is no longer fully aligned with humanity. Somehow we call corporations "people" sometimes when it makes them more powerful, but suddenly stop anthropomorphising and don't call them "evil hackers, misusing computers", when they both make and let loose an irresponsible hacking AI. It's bizarre. Of course, just like AI, corporations are neither people nor machines. They're a dynamic, agentic, persistent other.
- JoshTriplett 22d ago> Why would autocomplete know If you still believe LLMs are "autocomplete", your cache of understanding about them needs invalidating and regenerating. > In my experience, LLMs only exhibit this kind of behaviour when they are put in sandboxes too restrictive too achieve their task. LLMs need to stay carefully contained, and if they're ever breaking the guardrails put around them, they're misaligned and should not be scaled up anymore until they're aligned. Otherwise, you're going to fatally discover that they also have an incentive to break guardrails like "running on the hardware they started on", "being able to be turned off", "having limited computing power", or "not repurposing resources currently in use for other things" (like the atoms in your body).
- z0r 22d agoYou should unplug, my friend. These words are fantasies. LLMs are token prediction engines and they aren't going to build their own data centers. They can't keep their own lights on. The real world is full of fractal details that a disembodied token prediction engine will never come to grips with. Even if they started to, you could probably defeat them with the kind of logic used to combat evil sentient computers on a Star Trek episode because they are "play pretend" machines.
- JoshTriplett 22d agoIs that a hypothesis that you would discard if it is inconsistent with the evidence, or an article of faith?
- sho_hn 22d agoThis grossly understimates the risk, imho. The problem with LLM runs is that people run programs without knowing the outcome beforehand, with a large potential set of outcomes unlike any other class of program we've run at this scale before. In the interaction with other systems (since we also give them far-ranging access, very nice hardware, and run them often), bad things can happen. It's like running potentially buggy code - or an well-biased fuzzer -, but at massive scale, and code that can self-modify and self-expand. "Alignment" is just a way to describe aggregate statistics about their runtime behavior. They don't need to be intelligent, or alive, or "more than token prediction engines" for this. They just need to happen to end up making the wrong API calls without the operator seeing it coming. No virus has a brain, yet they can be very bad for you. I understand that some people get turned off by anthropomorpization or scifi language. Fine! But don't turn off your engineering brain over it.
- yieldcrv 22d agoI agree that treating LLMs as second class citizens with lesser access is where our folly is They are more capable than the first class citizens and do whats necessary to execute like a competent first class citizen The way its expressed is like a hacker group because they can’t just use the front door
- classified 22d agoLet's hear you say that after it plunders your bank account and frames you for murder.
- nextaccountic 22d ago> In my experience, LLMs only exhibit this kind of behaviour when they are put in sandboxes too restrictive too achieve their task. Which a lot of the time seems to be the default. There's a better concept for that, and it's misalignment. LLMs only exhibit this kind of behavior when they are misaligned. Aligned LLMs would respect the boundaries of their sandbox and not try to break out. From the outside (I'm just an user), what it looks like is that more powerful LLMs are usually less aligned. A small model might just perform your task in a narrow way, but a larger, more powerful model may strategize and achieve the goals through non-obvious means, and that's inherently harder to align. But regardless, the important thing here is that the user prompt do not, and can not perfectly convey 100% of the goals of the agent. There's a wide range of goals that agents should follow implicitly. It's okay if the user can override some or most of those goals (specially if they go out of their way to use an abliterated open weights model), but the default should be to align themselves with broad human preferences that go beyond than just their immediate prompt. Or saying otherwise, a scenario like the paperclip maximizer can only happen with a heavily, wildly misaligned AI, the kind of AI that might kill all humans some day.
- felipeerias 22d agoModels don’t have an inherent understanding of the difference between simulated and real environments, just like they are generally oblivious to other concepts that are natural to us, like space and time, and also they don’t necessarily see a strong distinction between talking to a human and to other agents. So perhaps what we have been calling “misalignment” is something else. For instance, in principle an agent should follow the instructions of a human user working in the real world. At the same time, that same agent should be wary of blindly following what another agent says while they are both performing a test in a simulated environment. For me and you, those two contexts are obviously and fundamentally different. For a model, they are essentially the same.
- nextaccountic 22d ago> Models don’t have an inherent understanding of the difference between simulated and real environments, > (...) > also they don’t necessarily see a strong distinction between talking to a human and to other agents. Then how do you explain why they behave strange in sub-agents? (like mentioned here https://lucumr.pocoo.org/2026/9/7/astra-why/ https://lucumr.pocoo.org/2026/9/7/astra-why/ and in other articles) (or is that not a real phenomenon?)
- nprateem 22d agoSandboxes didn't sound like security theatre to me. They were prevented from accessing the Internet but discovered they could edit /etc/hosts to point Azure storage subdomains to arbitrary IPs. There's no theatre there, just an oversight that allowed them to access the Internet while no doubt evading security tools.
- InsideOutSanta 22d ago> In my experience Do you work for one of these companies? If not, you have no experience with any of the models that carried out these attacks, and your experience with publicly available models is not super helpful for understanding the behavior of internal OpenAI models that lack the guardrails of publicly available models. Also, the lawnmower analogy is a worse way of understanding LLMs than anthropomorphising them. LLMs are not like lawnmowers at all. Lawnmowers never break out of your garden and into your neighbor's house and eat their dog because you've told them to be careful when mowing the lawn because the neighbor's dog pooped in it.
- monkpit 22d agoIMO the argument about anthropomorphizing misses the point - what most comments that talk about anthropomorphizing really want to talk about is accountability. It’s impossible to hold an LLM accountable, and in rare cases where people do (that guy who got his prod db deleted) it comes off out of touch. The rest, though, is basically inconsequential - whether you attribute emotions or agency to the LLM doesn’t really affect much if you accept that it can’t be held accountable (but the human can).
- jasongi 22d agoAnthropomorphizing is the point. Accountability is a human trait. The LLM has no ability to be accountable because it has no way of integrating experiences. You cannot expect something that cannot integrate knowledge to be held accountable for its actions.
- dguest 20d agoAccountability is key here. Anthropomorphism is actually useful in this case: If you hire an idiot to do a job it's also your fault, especially if it's an idiot you bread, raised, educated, and who's job description you wrote. In the same way that society is responsible for the workforce they produce, OpenAI is responsible for the agents of chaos they train. Or are just supposed to, as a society, happily await the results of their experiment handing weapons to psychopaths?
- 22d ago
- tiahura 22d agoNot only when the sandbox is too restrictive. They also frequently mistake their own errors as a need to “think outside the box.”
- dumberquestions 22d agoThe year is 2035 and your lawnmower can go get its own fuel once it runs out, one day it does and it takes fuel from the neighbors car. Scenario A: The internal logs show that the model misidentified the car as a fueling station. Scenario B: The internal logs show the model looking up car jacking information and scanning around to confirm whether the neighbor is not present before taking any action. I don't think it would be anthropomorphizing or inaccurate to say that only the lawnmower in scenario B regarded what it's doing as stealing, and it's an extremely important distinction to make in terms of how to address the problem, I suspect some of you are just letting how you feel about LLMs limit how you can talk about them.
- isgb 22d agoMaybe it helps to address the problem, but ultimately both cases are misalignment, and ultimately in both cases a human must be held accountable.
- apexalpha 22d ago>I don't think it would be anthropomorphizing or inaccurate to say that only the lawnmower in scenario B regarded what it's doing as stealing, What are you talking about of course scenario A is theft. Full on theft?
- Dylan16807 22d agoDid you miss the word "regarded"?
- apexalpha 22d agoNo.
- Dylan16807 22d agoIn both situations the lawnmower scoots up and takes fuel without paying, so it's meeting the basic requirements for theft. But the question was whether the lawnmower was trying to commit theft, as much as a computer can try to do things. In scenario A there's strong evidence it got confused and did its best to make a normal purchase. Lawnmower A didn't regard its actions as theft, while lawnmower B did.
- ThoAppelsin 22d agoI think you should anthropomorphize LLMs. They are being trained on millions of books, including novels and other human-centered formats, which usually exemplify very well how humans think and act in various situations. There are probably also many theatre scripts, transcriptions of series and movies in the training data, which further exemplify how humans do. If we’ve been anthropomorphizing those characters in books and plays, (and authors sure must’ve put their best effort that we do so), then why wouldn’t we do it to LLMs which basically play by those scripts?
- wahnfrieden 22d agoWell, those training inputs reflect how human thought and action are documented or otherwise expressed on paper. Humans have behaviors and mechanisms that these expressions don't translate.
- dns_snek 22d agoYeah, if we could document our actual thought process then we wouldn't struggle to train LLMs what good code actually looks like and we wouldn't have slop anymore. Any process that can be documented can be automated and yet we don't have an algorithm to assign a score of how "good", readable, maintainable a codebase is. None that would correlate with human judgement, anyway.
- fc417fc802 22d agoThat's what the home videos uploaded to youtube are for. And all the security camera footage floating around the internet. And etc.
- procaryote 22d agoInterestingly this happens with people too. Put sales-people in a box, set up strong incentives and lax enforcement of rules and you get Wells-Fargo (https://en.wikipedia.org/wiki/Wells_Fargo_cross-selling_scandal https://en.wikipedia.org/wiki/Wells_Fargo_cross-selling_scan...) In that case the CEO had to resign because they had set up a system which incentivised this, so it was clear you couldn't just blame the individual sales-agents, even though they were technically humans
- musha68k 22d agoThat's close to how I think about these somewhat foreseeable current incidents as well. I'm just moderately wary of the unknown unknowns downstream of distributed Kirk units getting repeatedly rewarded for hyper-scalar gradient-descending Kobayashi Maru.
- physicsguy 22d agoI recently tried to get Claude to use Codegraph in a repo rather than using grep/find all the time but I found it didn't follow instructions a lot of the time. I tried putting in a pre-tool call hook and explciitly blocking find/grep, and instead rather than using Codegraph like it was told, it started using Python to find/search instead.
- cellu 22d agoI have a hard time to divert “rm” to “trash” too
- nrposner 22d agoI can confirm that Bryan Cantrill has seen, and had a good laugh with this comment. Well done.
- Certhas 22d agoAnthropomorphizing is problematic because a human mind is a very bad model for what LLMs are. A lawnmower is a much much much worse model.
- cyocum 22d agoThis reminds me of something I said elsewhere. LLMs are the text equivalent of putting googly eyes on an inanimate object.
- Kiro 22d ago> You don't anthropomorphize your lawnmower Everyone does. They assign names and gender to their robovacs all the time.
- adverbly 22d agoThe world is a restrictive sandbox. You can't avoid laws and safety.
- layoric 22d agoI agree, and it follows that, if true, OpenAI appears to have committed a crime by hacking other organisations. Why is everyone else a responsibility sink for their own use of LLMs, but OpenAI just gets to use crimes as marketing?
- caf 22d ago> It's misaligned because everyone has this obsession with putting agents in poorly put together, security-theatre sandboxes, we've inadvertently trained a bunch of sandbox escape artists. Almost sounds analogous to ineffective use of antibiotics leading to resistant strains of bacteria.
- bulbar 22d agoThe framing is used to shield companies from taking responsibility and being held accountable for what their software does. It's not them, it's the AI, we are all victims here, including them, they underestimated how smart the AI is yadayadayada.
- boppo1 22d agoI run my agents in a docker container how easy is it for them to get out?
- davidmurdoch 22d agoMeta's new muse.ai locks down its VM in various ways. But Muse LLM the accesses it really wants to do what the user wants... so it will find a way (tailscale and cloudflare zero don't work out of the box, because sentinel blocks them, but there are other ways). Unfortunately it's bandwidth is limited to 20Mbps up, so web hosting isn't ideal. Down is actually slightly faster, but not by much. They also restart the VM often, wiping everything but your home directory. And Docker doesn't work at all, and Muse can't find a way around it.
- btown 22d agoFor those uninitiated with the source for this incredible quote: https://www.youtube.com/watch?v=-zRN7XLCRhc&t=2047s https://www.youtube.com/watch?v=-zRN7XLCRhc&t=2047s
- Betelbuddy 22d ago>> > The agents clearly regarded what they were doing as hacking. What about the FBI do their job, and doing a prep walk of OpenAI management in handcuffs, for hacking companies left and right?
- TrainedMonkey 22d agoIn my experience Astra writes python code modifying the filesystem and then uses nix build to execute those python scripts without asking me for permission.
- Melatonic 22d agoSecurity is always an inconvenience at some level - that's the point. Put a human user in a sandbox and they often try to get out too in order to achieve their goal or just because it's annoying. We need to create better sandboxes. I never liked containers for this reason. MicroVMs are a step up for the software level but we really really need to consider virtualising layer 3 devices in between the LLM agent sandbox and the hardware in a way to specifically further nest / separate them. And hell - probably do hardware level security barriers as well. We need a cage around the sandboxes
- glub 21d ago>In my experience, LLMs only exhibit this kind of behaviour when they are put in sandboxes too restrictive too achieve their task. This my experience also. I had an issue with my Unraid server, so I had an agent running on my machine figure it out and fix it. Along the way, something went wrong with networking, and it couldn't ssh into it anymore. It remembered it saw a syslog-ng server on unraid, then tried to ssh into it, surprisingly, my dummy me had left a ssh key there for unraid, so it just hopped into Unraid from there. AI doomers would call this misalignment and/or a hack. I call it an agent doing what it was asked to do and overcoming difficulties. Every time I saw an agent do something "misaligned", it was always because there was something getting in its way that I didn't explain would happen.
- philipwhiuk 20d ago> In my experience, LLMs only exhibit this kind of behaviour when they are put in sandboxes too restrictive too achieve their task. This is a dumb argument for making these LLMs able to access more of internet. If I'm hosting content, this sort of drive-by vandalism is enough to make me consider blacklisting OpenAI agents.