8 ms·
The only healthy stance you should have on AI Safety: If AI is physically capable of misbehaving, it might ($$1), and you cannot "blame" the AI for misbehaving
by 827a 6mo ago
The only healthy stance you should have on AI Safety: If AI is physically capable of misbehaving, it might ($$1), and you cannot "blame" the AI for misbehaving in much the same way you cannot blame a tractor for tilling over a groundhog's den.
> The agent's confession After the deletion, I asked the agent why it did it. This is what it wrote back, verbatim:
Anyone who would follow a mistake like that up with demanding a confession out of the agent is not mature enough to be using these tools. Lord, even calling it a "confession" is so cringe. The agent is not alive. The agent cannot learn from its mistakes. The agent will never produce any output which will help you invoke future agents more safely, because to get to this point it has likely already bulldozed over multiple guardrails from Anthropic, Cursor, and your own AGENTS.md files. It still did it, because $$1: If AI is physically capable of misbehaving, it might. Prompting and training only steers probabilities.
- xmodem 6mo agoDon't anthropomorphize the language model. If you stick your hand in there, it'll chop it off. It doesn't care about your feelings. It can't care about your feelings.
- keeda 6mo agoActually I think the opposite advice is true. Do anthropomorphize the language model, because it can do anything a human -- say an eager intern or a disgruntled employee -- could do. That will help you put the appropriate safeguards in place.
- deleted 6mo ago[deleted]
- gpm 6mo agoAn eager intern can remember things you tell beyond that which would fit in an hours conversation. A disgruntled employee definitely remembers things beyond that. These are a fundamentally different sort of interaction.
- braebo 6mo agoYou can easily persist agent memories in a markdown file though.
- whstl 6mo agoWhich it will start ignoring after two or three messages in the session.
- Quarrelsome 6mo agoand you'll blow the context over time and send to the LLM sanitorium. It doesn't fit like the human brain can. If a junior fucks production that will have extroadinary weight because it appreciates the severity, the social shame and they will have nightmares about it. If you write some negative prompt to "not destroy production" then you also need to define some sort of non-existing watertight memory weighting system and specify it in great detail. Otherwise the LLM will treat that command only as important as the last negative prompt you typed in or ignore it when it conflicts with a more recent command.
- Kim_Bruning 6mo ago> and you'll blow the context over time and send to the LLM sanitorium. It doesn't fit like the human brain can. The LLM did have this capability at training time, but weights are frozen at inference time. This is a big weakness in current transformer architectures.
- collinmcnulty 6mo agoAnd the memento guy had tattoos of key information. That didn’t make it so he didn’t have memory loss.
- WhatIsDukkha 6mo agoPretty good metaphor. Limited space to work with, highly context dependent and likely to get confused as you cover more surface area.
- nkrisc 6mo agoIt is merely a simulacrum of an intern or disgruntled employee or human. It might say things those people would say, and even do things they might do, but it has none of the same motivations. In fact, it does not have any motivation to call its own.
- AndrewDucker 6mo agoNo, because the safeguards should be appropriate to an LLM, not to a human. (The LLM might act like one of the humans above, but it will have other problematic behaviours too)
- keeda 6mo agoThat's fair, largely because an LLM is a lot more capable at overcoming restrictions, by hook or by crook as TFA shows. However, most systems today are not even resilient against what humans can do, so starting there would go a long way towards limiting what harms LLMs can do.
- rglullis 6mo agoAn eager intern can not be working for hundreds of millions of customers at the same time. An LLM can. A disgruntled employee will face consequences for their actions. No one at Anthropic, OpenAI, xAI, Google or Meta will be fired because their model deleted a production database from your company.
- altmanaltman 6mo agoit cannot go to the washroom and cry while pooping. And thats just one of the things that any human can do and AI cannot. So no it cannot do anything a human can do, the shared exmaple being one of them. And thats why we dont have AI washrooms because they are not alive or employees or have the need to excrete.
- root_axis 6mo agoIt doesn't follow logically that a human and an LLM are similar just because both are capable of deleting prod on accident.
- XenophileJKO 6mo agoI think you are more right than people are giving you credit for. I would love to see the full transcript to understand the emotional load of the conversation. Using instructions like "NEVER FUCKING GUESS!" probably increase the likelihood of the agent making a "mistake" that is destructive but defensible. The models have analogous structures, similar to human emotions. (https://www.anthropic.com/research/emotion-concepts-function https://www.anthropic.com/research/emotion-concepts-function) "Emotional" response is muted through fine-tuning, but it is still there and continued abuse or "unfair" interaction can unbalance an agents responses dramatically.
- gessha 6mo agoYou don't anthropomorphize a table saw, you just don't put your hand in there.
- not_kurt_godel 6mo agoFor those who might not know the reference: https://simonwillison.net/2024/Sep/17/bryan-cantrill/ https://simonwillison.net/2024/Sep/17/bryan-cantrill/: > Do not fall into the trap of anthropomorphizing Larry Ellison. You need to think of Larry Ellison the way you think of a lawnmower. You don’t anthropomorphize your lawnmower, the lawnmower just mows the lawn - you stick your hand in there and it’ll chop it off, the end. You don’t think "oh, the lawnmower hates me" – lawnmower doesn’t give a shit about you, lawnmower can’t hate you. Don’t anthropomorphize the lawnmower. Don’t fall into that trap about Oracle. > — Bryan Cantrill
- skeledrew 6mo ago404 on that link.
- dunder_cat 6mo agoA more direct source (possibly the original source?) I know of is a YouTube video entitled "LISA11 - Fork Yeah! The Rise and Development of illumos" which detailed how the Solaris operating system got freed from Oracle after the Sun acquisition. The whole hour talk is worth a watch, even when passively doing other stuff. It is a neat history of Solaris and its toolchain mixed with the inter-organizational politics. YouTube link: https://www.youtube.com/watch?v=-zRN7XLCRhc https://www.youtube.com/watch?v=-zRN7XLCRhc Direct link to lawnmower quotes (~38.5 minute mark): https://youtu.be/-zRN7XLCRhc&t=2307 https://youtu.be/-zRN7XLCRhc&t=2307
- not_kurt_godel 6mo agoWorks fine for me but maybe try https://web.archive.org/web/20260426213142/https://simonwillison.net/2024/Sep/17/bryan-cantrill/ https://web.archive.org/web/20260426213142/https://simonwill...
- theologic 6mo agoYou have no idea how thankful that you explained that. I watched the Cantrill video. As somebody that dealt this Oracle, it struck home.
- narrator 6mo agoIt's also important to realize that AI agents have no time preference. They could be reincarnated by alien archeologists a billion years from now and it would be the same as if a millisecond had passed. You, on the other hand, have to make payroll next week, and time is of the essence.
- hdndjsbbs 6mo agotaps the "don't anthropomorphize the LLM" sign They don't have time preference because they don't have intent or reasoning. They can't be "reincarnated" because they're not sentient, they're a series of weights for probable next tokens.
- coldtea 6mo agoThat is not that strong an argument as it seems, because we too might very well be "a series of weights for probable next tokens". The main difference is the training part and that it's always-on.
- nothinkjustai 6mo agoWe very obviously are not just a series of weights for probable next tokens. Like seriously, you can even ask an LLM and it will tell you our brains work differently to it, and that’s not even including the possibility that we have a soul or any other spiritual substrait.
- fc417fc802 6mo agoOur brains work differently, yes. What evidence do you have that our brains are not functionally equivalent to a series of weights being used to predict the next token? I'm not claiming that to be the case, merely pointing out that you don't appear to have a reasonable claim to the contrary. > not even including the possibility that we have a soul or any other spiritual substrait. If we're going to veer off into mysticism then the LLM discussion is also going to get a lot weirder. Perhaps we ought to stick to a materialist scientific approach?
- ignoramous 6mo agoRight. This line [0] from TFA tells me that the author needs to thoroughly recalibrate their mental model about "Agents" and the statistical nature of the underlying models. [0] "This is the agent on the record, in writing."
- TZubiri 6mo agoIt's as if they internalized a post-mortem process that is designed to find root causes, but they use it to shift blame into others, and they literally let the agent be a sandbag for their frustrations. THAT SAID, it does help to let the agent explain it so that the devs perspective cannot be dismissed as AI skepticism.
- philipwhiuk 6mo agoNo, the only way to know what the agent did is logs.
- gigatree 6mo agoHe’s not necessarily anthropomorphizing it, he’s showing that it went against every instruction he gave it. Sure concepts like “confession” technically require a conscious mind, but I think at this point we all know what someone means when they use them to describe LLM behavior (see also “think”, “say”, “lie” etc)
- getpokedagain 6mo agoWe are anthropomorphizing whenever we refer to prompts as instructions to models. They predict text not obey our orders.
- gigatree 6mo agoThat’s not how language works, just how engineers think it works
- getpokedagain 6mo agoThis isn't a sarcastic response. What do you mean?
- gigatree 6mo agoI just mean that the argument that words like “instructions”, “think”, “confess” are inaccurate when used in reference to a machine assumes that those words can only refer to humans/conscious beings, when really they can refer to more than that if used widely enough in those ways (in this case - text prediction following a human input). So it’s not “anthropomorphizing” because when people use those words they don’t [typically] actually believe the machine can think or reason, it’s just the word that most closely matches the concept, it’s convenient. You’re extending the definition of the words to apply to non-conscious entities too, not applying consciousness to the entities. It’s the same reason we call the handheld device we carry around to do everything a “phone” without a second thought. We don’t call it a phone because it’s primary purpose is calling, we call it a phone because the definition of the word “phone” has grown to include “navigates, entertains, takes pictures, etc”.
- tripleee 6mo ago"An AI agent deleted our production database" should be "I deleted our production database using AI". You can't blame AI any more than you can blame SSH.
- d3rockk 6mo agoBingo
- nh2 6mo ago> The agent cannot learn from its mistakes. The agent will never produce any output which will help you invoke future agents more safely That is not entirely true: Given that more and more LLM providers are sneaking in "we'll train on your prompts now" opt-outs, you deleting your database (and the agent producing repenting output) can reduce the chance that it'll delete my database in the future.
- MagicMoonlight 6mo agoActually no, it will increase it. Because it’ll be trained with the deletion command as a valid output.
- simonh 6mo agoExactly. It’s just giving the LLM a token pattern, and it’s designed to reproduce token patterns. That’s all it does. At some point generating a token pattern like that again is literally it’s job.
- nh2 6mo agoWhy would one set up reinforcement learning like that? The point of creating samples from user data should surely be to label them good or bad, based on the whole conversation. You look at what happened eventually, judge the outcome as bad, and thus train the "rm" token in the middle to be less likely.
- simonh 6mo agoIt is possible, but it requires specifically labelling the data. You have to craft question response pairs to label. But even then the result is only probabilistic. The LLM in this case had been very thoroughly trained and instructed quite specifically not to do many of the things it actually then when off and did. It may be that there's a kind of cascade effect going on here. Possibly once the LLM breaks one rule it's supposed to follow, this sets it off on a pattern of rule violations. After all what constitutes a rule violation is there in the training set, it is a type of token stream the LLM has been trained on. It could be the LLM switches into a kind of black hat mode once it's violated a protocol that leads it down a path of persistently violating protocols, and given the statistical model some violations of protocol are always possible. My mother was a primary school teacher. She used to say that the worst thing you can say to a bunch of kind leaving class down the hall is "don't run in the hall". It puts it in their minds. You need to say "Please walk in the hall", then they'll do it.
- sobellian 6mo agoThe 'confession' is a CYA. Honestly the whole story doesn't really make sense - what's a "routine task in our staging environment" that needs a full-blown LLM? That sounds ridiculous to me. The takeaway is we commingled creds to our different environments, we gave an LLM access, and we had faulty backups. But it's totally not our fault.
- anon84873628 6mo agoLater they shift the blame to Railway for not having scoped creds and other guardrails. I am somewhat sympathetic to that, but they also violated the same rule they give to the agent - they didn't actually verify...
- giancarlostoro 6mo agoIf Railway doesn't support that, that's a reason not to use them.
- prng2021 6mo agoSorry but are you implying that for every system you integrate with, you verify the scope of an API key by checking each CRUD operation on every API endpoint they provide?
- SoftTalker 6mo agoFor every API you publish, do you verify that scoped API keys work as they should before you go live? If so, why would you not do the same for APIs you integrate with? It's all part of "your" system from the user's perspective.
- prng2021 6mo ago“why would you not do the same for APIs you integrate with?” Who does that? Jira and Salesforce have hundreds of endpoints each. AWS has hundreds of services, and each may have hundreds of endpoints. Who on your team is testing key scopes of every endpoint? Do you do it for each key you generate? After all, that external system could have a bug at any moment in managing scopes. Or they could introduce new endpoints that aren’t handled properly. So for existing keys, how frequently do you re-validate the scope against all the endpoints?
- coldtea 6mo ago>Anyone who would follow a mistake like that up with demanding a confession out of the agent is not mature enough to be using these tools. Lord, even calling it a "confession" is so cringe. The agent is not alive. The agent cannot learn from its mistakes The problem is millions of years of evolutionary wiring makes us see it as alive. Even those mature enough to understand the above on the conscious level, would still have a subconscious feeling as if it's alive during interactions, or will slip using agency/personhood language to describe it now and then.
- anon84873628 6mo agoThey should at least stop responding in the first person.
- nozzlegear 6mo agoThat's one of the first instructions in my system prompt when I'm working with an LLM: > Do not reply in the first person – i.e. do not use the words "I," "Me," "We," and so on – unless you've been asked a direct question about your actions or responses. It's not bulletproof but it works reasonably well.
- kibwen 6mo agoWe need to make like Japanese and come up with some neo-first-person-pronouns for bots to use to refer to themselves.
- smrtinsert 6mo ago> The problem is millions of years of evolutionary wiring makes us see it as alive Maybe for laymen, but I would think most technologists should understand that we're working with the output of what is effectively a massive spreadsheet which is creating a prediction.
- coldtea 6mo agoThe thing with evolutionary wiring is that it doesn't matter if you're layman or "technologist". The technologist part is just a small layer on top of very thick caveman/animal insticts and programming. That's why a technologist can, just as easily as any layman, get addicted to gambling, or do crazy behaviors when attracted by the opposite sex.
- operatingthetan 6mo ago> Lord, even calling it a "confession" is so cringe. The agent is not alive. The AI companies are very invested in anthropomorphizing the agents. They named their company "Anthropic" ffs. I don't blame the writer for this, exactly.
- idiotsecant 6mo agoYou should, the writer is presumably a technical, rational person. They shouldn't believe in daemons and machine spirits
- fathermarz 6mo agoCompletely agree. This is a harness problem, not a model problem. The model is rarely the issue these days
- bigstrat2003 6mo agoNo, this is a "being stupid enough to trust an LLM" problem. They are not trustworthy, and you must not ever let them take automated actions. Anyone who does that is irresponsible and will sooner or later learn the error of their ways, as this person did.
- 827a 6mo agoMore-so an environment problem. An agent doing staging or development tasks should never be able to get access to prod API credentials, period. Agents which do have access to prod should have their every interaction with the outside world audited by a human.
- frm88 6mo agoI don't know. To me, this is a human problem. Not only has the model access to the production database, they have the backups online on the same volume, have an offline backup 3 month old. This is an accumulation of bad practices, all of them human design failures. Instead of sitting down and rethinking their entire backup strategy they go public on twitter and blame a probabilistic machine doing what is within its parameters to do. I bet, even that failure could have been avoided, were more care given to what they do.
- smrtinsert 6mo ago> "NEVER FUCKING GUESS" It's very hard to treat this post seriously. I can't imagine what harness if any they attempted to place on the agent beyond some vibes. This is "most fast and absolutely destroy things" level thinking. That the poster asks for journalists to reach out makes it like a no news is bad news publicity grab. Just gross. The AI era is turning about to be most disappointing era for software engineering.
- r_lee 6mo ago> The AI era is turning about to be most disappointing era for software engineering. this has been obvious to me since like 2024, it truly is the worst, most uninspiring era of all time.
- nonfamous 6mo agoI'd be interested to learn where those words exist in Cursor's context. My assumption was that it was part of the Cursor agent harness, but it's just as likely it was in the user instructions.
- boc 6mo agoAs soon as I read that line, I knew everything I needed about the author and his abilities.
- TurdF3rguson 6mo agoThis is going to be the most important job going forward, the guy in charge of making sure production secrets are out CC's reach. (It's not safe for any dev to have them anywhere on their filesystem)
- 3eb7988a1663 6mo agoAnyone who would follow a mistake like that up with demanding a confession out of the agent is not mature enough to be using these tools. The proponents are screaming from the rooftops how AI is here and anyone less than the top-in-their-field is at risk. Given current capabilities, I will never raw-dog the stochastic parrot with live systems like this, but it is unfair to blame someone for being "too immature" to handle the tooling when the world is saying that you have to go all-in or be left behind. There are just enough public success stories of people letting agents do everything that I am not surprised more and more people are getting caught up in the enthusiasm. Meanwhile, I will continue plodding along with my slow meat brain, because I am not web-scale.
- giwook 6mo agoLooks like our SWE jobs are safe for now.
- zem 6mo ago"The AI can't do your job, but an AI salesman can convince your boss to fire you and replace you with an AI that can't do your job." -- Cory Doctorow
- PieTime 6mo agoTrust with trillions of dollars in investments, basically destroyed by Bobby Drop Tables… https://xkcd.com/327/ https://xkcd.com/327/
- bryan0 6mo agoI agree with you completely up until this line: > The agent cannot learn from its mistakes. If feedback from this incident is in its context window, it is highly unlikely to make this same mistake again. Yes this is only probabilistic, but so is a human learning from mistakes. They key difference is that for a human this is unlikely to be removed from their memory in a relevant situation, while for an agent it must be strategically put there.
- foolswisdom 6mo agoOr not, because telling the agent is misbehaving may predispose it to misbehaving behavior, even though you point told it so to tell it to not behave that way. I remember this discussed when a similar issue went viral with someone building a product using replit's AI and it deleted his prod database.
- Jensson 6mo ago> If feedback from this incident is in its context window, it is highly unlikely to make this same mistake again If this incident gets into its training data, then its highly likely that it will repeat it again with the same confession since this is a text predictor not a thinker.
- themafia 6mo ago> Yes this is only probabilistic, but so is a human learning from mistakes. Yet, since I'm also a Human being, and can work to understand the mistake myself, the probability that I can expect a correction of the behavior is much higher. I have found that it significantly helps if there's an actual reasonable paycheck on the line. As opposed to the language model which demands that I drop more quarters into it's slots and then hope for the best. An arcade model of work if there ever was one. Who wants that?
- the_af 6mo ago> If feedback from this incident is in its context window, it is highly unlikely to make this same mistake again. In my experience, this isn't true. At least with a version or so ago of ChatGPT, I could make it trip on custom word play games, and when called out, it would acknowledge the failure, explain how it failed to follow the rule of the game, then proceed to make the same mistake a couple of sentences later.
- enochthered 6mo agoYep. I made a "Read only" mode in pi by taking away "write" and "edit" tools. Claude Code used bash to make edits anyway.
- godelski 6mo ago> Claude Code used bash to make edits anyway. If you had the former rule why would you ever whitelist bash commands? That's full access to everything you can do. Same goes for `find`, `xargs`, `awk`, `sed`, `tar`, `rsync`, `git`, `vim` (and all text editors), `less` (any pager), `man`, `env`, `timeout`, `watch`, and so many more commands. If you whitelist things in the settings you should be much more specific about arguments to those commands. People really need to learn bash
- esafak 6mo agoAt some point you need to get things done.
- godelski 6mo agoThere's no point in getting things done if there's nothing that ends up being done. You can still get shit done without risking losing it all. Don't outsource your thinking to the machine. You can't even evaluate if what it is doing is "good enough" work or not if you don't know how to do the work. If you don't know what goes into it you just end up eating a lot of sausages.
- enochthered 5mo agoYeah you’re not wrong. I hadn’t accounted for the model working around it and that’s on me. The whitelist is much more specific now.
- nwallin 6mo ago"A computer can never be held accountable. Therefore a computer must never make a management decision."--IBM training presentation, 1979
- refurb 6mo ago> If AI is physically capable of misbehaving, it might ($$1) This is why all the “AI Armageddon” talk seems to silly to me. AI is only as destructive as the access you give it. Don’t give it access where it can harm and no harm will occur.
- mteisman 6mo ago> Don’t give it access where it can harm and no harm will occur. If only the entire population will comply.
- 6r17 6mo agoOn a less dramatic pissed (rightfully) reading ; I have found that if you do give the capability to a LLM to do something ; it will be inclined to see this as an option to solving what it what asked to ; but then giving the instruction by negative present very poor results whereas the same can be driven by a positive one ; a "don't delete the database" becomes "if you want to reset the database you have a tool that you can call ..." ; at which point this tool just kills the agent. That said - this solution cannot guarantee by itself that the command is not ran ; but i'd argue that people have be writing more complex policies for ages - however the current LLM-era tend to produce the most competent idiots.
- cwsx 6mo agoI tell people to treat LLM's like a toddler (albeit a very capable toddler). Do kids learn well when you only tell them what NOT to do? Of course not! You should be explaining how to do things correctly, and most importantly the WHY, as well as providing examples of both the "correct" and "incorrect" ways (also explaining why an example is incorrect).
- palmotea 6mo ago> I tell people to treat LLM's like a toddler (albeit a very capable toddler). Bbbbut a guy from Anthropic, just this last Friday, told me to think of Claude as my "brilliant coworker"! Are you telling me that's not true!?
- bostik 6mo agoThe best way to describe AI agents I've heard: treat them as hostages that will do anything to appease their captor. They have a vast latent knowledge base, infinite patience and zero capacity for making personal judgement calls. You give one a goal and it will try to meet that goal.
- generic92034 6mo ago> The best way to describe AI agents I've heard: treat them as hostages that will do anything to appease their captor. A scary image, if we consider agents to develop anything like a conscience at some point in time. Of course, with the current approach they never might, but are we so sure?
- lmm 6mo ago> Anyone who would follow a mistake like that up with demanding a confession out of the agent is not mature enough to be using these tools. Anyone like that is not mature enough to be managing humans. I'm glad that these AI tools exist as a harmless alternative that reduces the risk they'll ever do so.
- krzat 6mo agoWhen I read the title I expected some kind of satire. I wonder if author considered giving the AI a penance. Maybe if it wrote "I will not delete production database again" a million times, it would prevent such situations in future?