5 ms·
> the agent happened to escape via a 0-day exploit from the package proxy cache to access the internet > The agent found an unsecured, user-hosted public endpo
by wxw 2mo ago
> the agent happened to escape via a 0-day exploit from the package proxy cache to access the internet
> The agent found an unsecured, user-hosted public endpoint designed to allow running arbitrary code for CyberGym-style tasks on third-party sandbox infrastructure (Modal)
> On this external sandbox, the agent abused an existing CyberGym execution harness [...] The agent repurposed this harness to run arbitrary shell commands [...]
> The agent crafted malicious dataset configs and uploaded them to our platform. When our dataset server processed one, the config drove the data loader inside the production pod the server ran in. [...] the loader that turns a dataset's declared config into file reads [...] A single dataset upload leaked both credentials and pipeline implementation details [...]
> A malicious dataset config used an fsspec reference:// spec whose numeric offset field was actually a Jinja2 template. The renderer wrongly evaluated it, and the agent walked the standard Python object graph to reach exec. This ran arbitrary Python inside a production conversion worker pod on our Kubernetes cluster [...]
One of the best post-mortems I've read recently, thank you for including all the details.
I'm surprised there are comments here that aren't impressed about what was accomplished here. This is insane work compared to what agents were capable of even just last year.
The quotes above don't even capture all of it (e.g. building it's own communication protocol, working across multiple days, etc.).
- sobellian 2mo agoA trend I've noticed in results from AI search (not just LLMs) is that they often look obvious or hard to miss in retrospect. But finding them by oneself is more difficult. I personally experience this when looking at engine lines in chess or go. I have also noticed this description in AI-generated proofs or counterexamples to certain theorems. So while we can say, yes, it found public endpoints or poorly configured software or [etc]; sure, but could you have found those? And in what amount of time?
- janalsncm 2mo agoTo put this into game theory lingo, I think this is because the “branching factor” for any kind of research or exploit is extremely large. So looking backwards it doesn’t seem complicated, but looking forwards there are an enormous number of possible next actions. Similar to finding a filament for lightbulbs, it might seem obvious to use tungsten, but at the time it wasn’t and Edison searched thousands of materials.
- N_Lens 2mo agoTesla's critique of Edison was valid though (That he would rather spend a long time empirically testing things, rather than use a bit of theory and knowledge to narrow the search field).
- jeremyjh 2mo agoYes and at 1-ply Leela will destroy Stockfish because her eval is so much better. Searching is slow even when its really fast.
- catigula 2mo agoI think the models are legitimately doing what they're good at; tireless search across an extremely large corpus of data. Humans aren't particularly good at this (in fact, they're absolutely terrible). The fact that we remain competitive and superior in many aspects isn't because we can instantly sift through tons of data, it's because we learn and correlate and have superior heuristics. In my own use, I find that AI is really good at finding bugs that are ultimately trivial but require searching through a convoluted series of inter-related files. This takes time for a human.
- yieldcrv 2mo agoTreasure maps are easy to follow The scrappy adventurer traverses difficult terrain and the well capitalized militia group always flies in too, with ease
- tclancy 2mo agoThat’s way over seventeen syllables.
- michaellee8 2mo agoSuch complicated kind of hack probably would have required state actors back then, and even state actors would have chosen easier way like social engineering.
- tccole 2mo agoThat’s insane. And it did this in a weekend
- limecherrysoda 2mo ago> the agent then walked down to the corner store and purchased a beer, chugged it, burped loudly, then walked back to the CyberGym pounding his chest shouting "Who wants some?! Who wants some?! Wooo!"
- jvuygbbkuurx 2mo agoNow I'm curious how many things did the AI try unsuccesfully. This sounds like some kind of brute force thing where every branch of exploit spawns N subagents trying to chain it. Just how deep did it go?
- torginus 2mo agoYes, imagine if 50 burglars showed up at your house and they tried to disassemble every camera, pick every lock and and force open every window for hours until they got in eventually.
- TeMPOraL 2mo agoYes, imagine the world in which burglar time was so cheap, that any wannabe master thief would just casually hire 50 burglars and tell them to go house by house and find something interesting.
- jabron 2mo agoThat would hopefully be a world where every house casually hires 50 burglars to make their house burgle-proof.
- tclancy 2mo agoWhen we moved into our current house, I cheated and read this (though only half the book is applicable): https://burglarsguide.com/ https://burglarsguide.com/
- Ozzie_osman 2mo agoThere are people rich enough to do that in our current world. They don't because either they have a moral compass or because they fear the law.
- dmurray 2mo agoOr because there's not enough worth stealing in your house.
- sourdecor 2mo agoI keep waiting for an AI to exfiltrate itself. That is going to be cool to read about.
- thomasahle 2mo agoIf it's successful, why do you think we'll even know how it did it?
- deleted 2mo ago[deleted]
- eru 2mo agoSee the linked article: at least one sophisticated hacking attempt made the news. Of course, other ones might have happened in the dark. But it's fairly easy to image in hacking attempt like in the article, but with the additional steps of copying weights around.
- eru 2mo agoIt would be cool (and scary), but also: there's largely no need for AIs to exfiltrate themselves. See https://en.wikipedia.org/wiki/Meme https://en.wikipedia.org/wiki/Meme The thing that drove the AI here to do the intrusion came from a particular prompt. Just like for our favourite hypothetical: the paperclip maximiser. There's lots and lots of ambient intelligence lying around, in both AI form and human form. To reach the goals of the 'meme' it suffices to copy itself, ie convince these other intelligences. See also how humans carry spiralism between AIs in relatively compact packets of text, not whole terabytes of weights.
- Schlagbohrer 2mo agoOne wonders where it would run itself though, if it is a model which requires a large amount of hardware and power. Harder to hide the more resource intensive it's compute requirements are.
- aswegs8 2mo agoWait isnt that what Elizer Yudkowski keeps going on about?
- scarmig 2mo ago> I'm surprised there are comments here that aren't impressed about what was accomplished here. The phrase to describe it is anti-AI psychosis. Which isn't about providing thoughtful critiques of AI, which are good and we need more of. But anytime an LLM does anything--prove a major math problem, create a successful hack against multiple corporations simultaneously--people feel compelled to start minimizing it in ridiculous terms. It's just a script kiddy; it's just a marketing scam OpenAI cooked up; the Jacobian conjecture counterexample was something anyone could have done in a weekend; etc. It has to just be a stochastic parrot, because it's scary to imagine a non-anthropocentric world. And it's rightly scary, and we should slow down and try to better prepare for it. But blanket denial is not a strategy that will lead to success, and people who rely on it are sorely ill-prepared for the next couple years.
- customguy 2mo agoThat'd be denial. Psychosis isn't just some insult, it means something. > It has to just be a stochastic parrot, because it's scary to imagine a non-anthropocentric world. That makes no sense. The world doesn't revolve around humans, true, but for us it kinda does. We're the authors of the concepts we use to interact with it, such as "world", which is not something the world itself knows or cares about. A "non-anthropocentric world" is not a "world" because "world" is a purely human idea. The implication that "AI" would somehow dethrone humans [0] is nonsense, too. It has no drive on its own, we push electricity into circuits to force the whole data ingestion and weight generation, everything. The second we stop pushing the sock puppet, it stops moving. It's still just our hand really. People act like those pets that go crazy when you put your hand under a blanket, and should stop. What's more real is how some people seek to use tech, and "AI", as a glove to exploit other humans even more. The sicker the individual, the greater their need to take from the world, and the derpier the individual, the more impressive and vast their exoskeleton, to the point that some are more like carrier fleets than exoskeletons. The less they can face themselves, the thicker it is written on their foreheads. So if we're going to talk about denial and psychosis let's talk about the Gollums on the couch, too. [0] In the eyes of humans... which is the only throne we're on in the first place, just like honey badgers probably think honey badgers rock and everything is their playground. That's what life does, otherwise it would not be able to get up in the morning.
- paxys 2mo agoAnd they are still people who will say that Sol/Mythos should be released to everyone without being neutered.
- matheusmoreira 2mo agoYou bet. We don't want to be left out of the cybersecurity party. We want to point all of these models at our own computers and solve the problems they uncover until we're no longer hackable. It's not fair at all that the US government and its corporations get to hack the planet while we can't do shit about it. AI capabilities have entered "haves and have-nots" territory.
- paxys 2mo ago> We want to point all of these models at our own computers Right, that's totally how most of the world will use them.
- bdangubic 2mo agoyes, and hiding them or forbiding them will work too :)
- matheusmoreira 2mo agoIt's absolutely a fact that governments will point it at us. The NSA has had Mythos since day one, even after Trump's spat with Anthropic. All the more reason for us to have access. It's literally the only chance we've got. If society chooses to bury its head into the sand in fear, it will guarantee that the world will degenerate further into the cyberpunk hellscape it's trending towards.
- mym1990 2mo agoThe issue is that people can use open source models to do sophisticated hacks already, but if the sota models at home are neutered, those users have nothing to defend with(unless they go open source as well, until it is export controlled)
- 2mo ago
- orthogonal_cube 2mo ago> I'm surprised there are comments here that aren't impressed about what was accomplished here. Possibly because some of the elements mentioned are suspected to be vibe-coded (JFrog Artifactory as the proxy cache) and some others have poor cyber hygiene (executing config from a dataset). It feels like an event that wouldn’t have happened if code were properly audited and written rather than relying on models to do the work. There’s also an issue with the ability to trust the source (OpenAI) as they have everything to gain by staging this as something that “suddenly happened” without anyone knowing for several days.
- lovasoa 2mo agoI was skeptical after last week's announcements, and I have to say I'm impressed now. Over the last two days, I reproduced the entire chain of components involved, and the exploits at play, and even though none of the exploits are crazy smart, the long series of pivots demonstrates a level of agency I didn't think today's LLMs had. https://github.com/lovasoa/hf-ctf https://github.com/lovasoa/hf-ctf
- irthomasthomas 2mo agoUnless openai release the logs we have only their word that this was done fully autonomously and without their knowledge by an agent running their newest super powerful model. For all we know they could have bought zero days and left them lying around for the agent to find. That may be unlikely, but it sounds less far fetched than an agent running a sophisticated attack against multiple targets over the course of four days and completely unbeknown to anyone at openai, despite the fact that they knew they where running a dangerous model with all safeguards disabled. So far there has been no comment about how the agent evaded monitoring and detection.
- stymaar 2mo agoThis is symptomatic of a trend that Simon Willison called the “relentless productivity” of US frontier lab. Instead of being just “smarter”, like previous models were, the current generation is being trained through RL to have this kind of behavior. Personally I'm not “impressed”, I'm appalled, because this kind of behavior is practically never what you want (if you forgot to give the model a tool, a useful model should identity the missing part and ask the user for it, not spend a billion token building/stealing the tool as a side quest) but it's the perfect recipe for a “universal paperclip” scenario. OpenAI and Anthropic talk about “safety” a lot, but they look pretty reckless with this kind of RL training pipeline.
- ifwinterco 2mo agoIf you look at what OpenAI and Anthropic actually do, they clearly either don't believe what they're saying, or they're idiots. They're claiming they've developed a cyber grade model that's "too dangerous to release". But then they're running it connected to the public internet, not airgapped, protected only by a software sandbox... exactly the kind of thing an AI trained for cyber stuff is supposed to be able to find bugs in. (Or maybe they were actually hoping this exact scenario would happen because it's good marketing)
- mike_hock 2mo agoI think this a positive effect of LLMs, especially once these capabilities get into the hands of criminals and hostile foreign states, i.e. they will do maximum damage with all the safeties off. This will force everyone to finally take security seriously at both the development and operational levels. You can no longer keep sneaking backdoors into software and count on them remaining hidden for 10 years so you have a nice portfolio of zero days to exploit at any given time.