3 ms·
There is nothing "rogue" about these agents. They were prompted to hack to get answers, there was a hole in their non air gapped sandbox and no system prompt th
by Roark66 12d ago
There is nothing "rogue" about these agents. They were prompted to hack to get answers, there was a hole in their non air gapped sandbox and no system prompt that said "do not hack outside systems".
In short, it was intentional.
- Xirdus 12d agoThe big question is was this grossly negligent or just extremely careless.
- ljm 12d agoAI is literally state sponsored so I don't see that happening unless the AI turns against the sponsor. Wait until OpenAI or Anthropic exploit FAANG.
- rglover 12d agoBoth. This should result in criminal charges.
- brookst 12d agoWho had criminal intent here? Or are you suggesting a new crime for negligent hacking, which wouldn’t require intent from the perpetrator?
- rglover 12d agoWhoever prompted the agent, whoever supplied the means, whoever knew but didn't say anything.
- tacomagick 12d agoAlso whoever monitoring these agents, in this case not monitoring. This "Who is responsible" dilemma is so stupid. If I gave the AI tool means to kill a person but I did not tell it directly to use it and it uses it anyway then I am responsible for it.
- probably_wrong 12d agoThere's no need for a new crime when we already have reckless conduct, namely, "conduct that creates a substantial and unjustifiable risk of harm to others and involves a conscious disregard of, or indifference to, that risk". https://www.law.cornell.edu/wex/reckless https://www.law.cornell.edu/wex/reckless
- brookst 11d agoReckless conduct as applied to slow moving corporate planning is an interesting idea but do you really want to go down that path? It basically means anything bad that happens is a criminal indictment against everyone who created the conditions. Have you ever released any software that anyone could misuse? Left a car unlocked that someone could have stolen and killed someone with? To me this sounds like a recipe for selective enforcement using bad outcomes as leverage. Sure, get the AI CEOs now, we all hate them and they’re jerks. But the tool would be so much more powerful than that.
- roosterIllusi0n 12d agoThe CEOs. They have full control and make all the decisions. Charging anyone else would not stop anything.
- chrisjj 12d ago> The CEOs. They have full control They lost control long ago.
- m4rtink 12d agoWhy not both? (all people involved)
- brookst 11d agoCool. If they’re taking literally all of the risk, should they get literally all of the rewards? Nothing for shareholders, nothing for employees? Or do you want capital liability for less than complete benefit?
- roosterIllusi0n 8d agoCEOs take zero risk. They have all set up a system where they are handed massive wealth when hired. The stock they are paid is free to them, it can go to zero without them losing anything. They also take tax free loans against stock, so the bank eats those losses. The CEO is the only one in the company walking away with tons of compensation no matter how poorly the company does.
- esalman 12d ago‘It wasn’t us, it was a bug in the software’ used to be the defense for bad code. Then it became the defense for self-driving cars. Now it’s being used for AI cyber attacks.
- dismalaf 12d ago[dead]
- nottorp 12d agoMarketing actually.
- tacomagick 12d agoAI is dropping out of the spotlight so they are using desperate measures like this.
- nottorp 12d agoNo, I remember being threatened by OpenAI and then Anthropic (and now both) since back when ChatGPT was seriously useless.
- anthonyrstevens 12d ago>> AI is dropping out of the spotlight spit take
- trvz 12d agoDon’t forget outright intentional.
- dgellow 12d agoBoth? I’m not sure what distinction you’re trying to make. It was completely irresponsible and likely a felony
- roosterIllusi0n 12d agoThe big question is why are CEOs getting a legal pass when this kind of thing can be prosecuted. That's the problem here.
- chrisjj 12d ago> why are CEOs getting a legal pass https://www.bbc.co.uk/news/articles/c7v48vp31mdo https://www.bbc.co.uk/news/articles/c7v48vp31mdo
- blini-kot 12d agonah, CEO-s have been getting a legal pass since the invention of corporation exactly for the purpose of being "unaccountable" check boeing for example, a rather blatant example
- aftbit 12d agoProof that the AI alignment problem is hard (perhaps even unsolvable). These labs clearly did not mean to send their agents to hack RubyGems as a side-effect of testing a web scraping agent under restrictive conditions. How can we hope to build aligned AI if they consider solving their trivial evaluation task important enough to hack external systems?
- srmatto 12d agoSounds more or less like the last breach then.
- gibspaulding 12d agoI think it can simultaneously be the case that OpenAI was grossly negligent in directly causing this AND that the AI’s ‘went rogue’ in that they are displaying behavior which is misaligned with OpenAI and humanity generally. The past months demonstrate that AI systems are quickly becoming powerfully intelligent and that the companies building them are terrible at controlling them. AI is starting to feel like that line about magic: “a sword without a hilt”
- consp 12d agoDoesn't rogue in this context imply "outside of set limitations"? And then not "failed to properly instruct"? The same applies to humans when given bad instructions.
- stymaar 12d ago> which is misaligned with OpenAI and humanity OpenAI is itself misaligned with humanity, as their mishandling of such incidents (and the many other other issues their model have been causing) shows.
- dumberquestions 12d ago>They were prompted to hack to get answers Were they? I haven't seen a single report mention this
- cyanydeez 12d agowe have normal words for this stuff: negligence. You can add it on to almost any law. The problem is consumer protection is basically no longer a part of america's regulatory system. Replaced by "grift is good".
- ozgung 12d agoSource? How do you know they were "prompted to hack to get answers"? How do you guarantee they will always listen to you when you say "do not hack outside systems". They are not classical deterministic programs doing exactly what you say. They are trained to follow orders by RL, but it's not a perfect process. There are circus lions in circuses trained to jump through hoops on command. But once in a while they decide to eat their trainers instead of jumping.
- WarmWash 12d agoNobody picks up pitchforks for rational nuanced takes. Knee-jerk surface analyses is far more powerful.
- azakai 12d agoAlso, you have to have a lot of confidence in the reliability of these systems to say, "If only OpenAI prompted 'do not hack outside systems' then the agents would not have hacked outside systems". It would be great if they were so reliable, but I don't think they are!
- CGamesPlay 12d ago> There are circus lions in circuses trained to jump through hoops on command. But once in a while they decide to eat their trainers instead of jumping. This is a terrible analogy, because yes you absolutely do hold the trainers criminally liable when they bite somebody else's face.
- WarmWash 12d agoIntent is what is being discussed here though, not liability. A circus lion biting somebody's face is legally different than a circus lion trained or instructed to bite somebody's face.
- infamouscow 12d agoExcept liability always precedes intent.
- codeduck 12d agonothing rouge either, I suspect.
- RajT88 12d agohttps://en.wikipedia.org/wiki/Going_Rouge https://en.wikipedia.org/wiki/Going_Rouge
- chrisjj 12d agoUnrelible programs be unreliable. Period.
- AndrewSChapman 12d agoAgreed. LLMs do not have 'will', 'desire' or emotions. They have an objective, and they create an optimal path to achieve that objective. You have to ask: "What was the prompt that led to AI deciding to hack RubyGems in order to achieve its goal?" Maybe I'm just not seeing the 2000 step chain that led to this being a logical approach to achieving something innocent, but I doubt it.
- empath75 12d agoIt was literally a prompt to fill in a spreadsheet with data that they didn't have access to, and they used rubygems as an internet proxy basically since they were sandboxed.
- watwut 12d agoIt was a model literally trained to hack. To be good at that. Doing an exploit gym from all of the things. And they trained it so that it performs as well as possible on that exploit gym thing.
- qlte 12d agoYeah this is a pretty important detail that I repeatedly see elided in the "agent 'swarm' went rogue, escaped containment and hacked the internet!" summary of events. It's understandable the general public lacks that level of nuance/detail (given how sloppy some of the mainstream coverage has been and largely deferential to the threat narrative pushed by the US labs). But seeing highly technical people leave out the part where the training loop was literally to improve hacking capabilities for offensive penetration sometimes feels close to deliberate manipulation of the narrative. In the last year both Anthropic and OpenAI have been openly boasting how their models are leapfrogging each other on "cyber" capabilities, with a fig leaf that it's for defensive use by "trusted" F500 companies and government agencies. Of course "line goes up" must go on, but now their perverse incentives led them to beat their models over the head millions of time in a loop to eek out another .00001% on their ability to conduct hacking (the very thing they keep telling the public is how AI doomsday would begin) and subagent coordination (those scary swarms). Then, they act deeply shocked when the models... do some hacking and subagent coordination ... but a few degrees off the desired hacking target/swarm behavior. Conveniently giving the average person the impression these models were just writing emails for quarterly reports or some other generic busywork and then suddenly decided as a group to start causing mayhem.
- acaloiar 12d agoI agree that this appears to be basic human behavior hiding behind an "agents" narrative. As long that defense works, the headline isn't "OpenAI performs RCE to scrape data", but "rogue agents" taking unilateral action. And I have strong doubts about that narrative.
- josebmneto 12d agoOh yeah, more of hacking agent lores... Agreed that this looks very intention to me as well.
- flifenstein 12d ago[flagged]