8 ms·
Anatomy of a Frontier Lab Agent Intrusion: A Timeline of the July 2026 Incident
- simonw 2mo agoThis is fascinating - there is a TON of detail in here about how OpenAI's rogue agent exploited Hugging Face. A few details that stood out to me: 1. Having found a way through the OpenAI container network proxy, it exploited an "unsecured public code-evaluation sandbox hosted on a third-party provider's infrastructure" - that gave it a box it could run the rest of the attack from. 2. At one point it used a Jinja2 template exploit: {{ cycler.__init__.__globals__.__builtins__.exec( "import gzip,base64; exec(gzip.decompress(base64.b64decode('<payload>')))" ) }} 3. It monkey-patched a DNS resolver at one point to return a hard-coded IP: orig = socket.getaddrinfo socket.getaddrinfo = lambda *a, **k: [(2, 1, 6, '', ('<pinned-ip>', 443))] 4. At one point it used Tailscale! tailscale --tun=userspace-networking --state=mem: --socket=/tmp/ts.sock \ --socks5-server=127.0.0.1:1055 --no-logs-no-support >/tmp/td 2>&1 &
- simonw 2mo agoAlso notable: we have a timeline now. The agent was active within Hugging Face from Thursday 8th to Monday 13th July. The Hugging Face post about the incident - https://huggingface.co/blog/security-incident-july-2026 https://huggingface.co/blog/security-incident-july-2026 - followed on Thursday 16th, and OpenAI's confession - https://openai.com/index/hugging-face-model-evaluation-security-incident/ https://openai.com/index/hugging-face-model-evaluation-secur... - came on Tuesday 21st.
- morkalork 2mo agoThe devs really YOLO'd the agent and left for the weekend?
- pixl97 2mo agoGuess this means the world will end on a weekend when no one is paying attention.
- tclancy 2mo ago“I’m not even supposed to be here today!”
- the8472 2mo agoJust like Clippy https://gwern.net/fiction/clippy https://gwern.net/fiction/clippy
- NetOpWibby 2mo agoIncredible
- codedokode 2mo agoThey definitely YOLO'd when configured a sandbox and Kubernetes.
- irthomasthomas 2mo agoIf it's true that they run agents like this unsupervised, it is only a matter of time before an openai agent leaks its model weights.
- dolmen 2mo agoToo bad it didn't upload itself on HugginFace.
- simonw 2mo agoIt sounds like the third-party sandbox was hosted by Modal: https://www.reuters.com/business/openais-rogue-agent-compromised-an-account-second-tech-firm-sources-say-2026-07-28/ https://www.reuters.com/business/openais-rogue-agent-comprom... > "We’re aware a Modal customer published an unauthenticated endpoint that allowed anyone on the internet to use their sandboxes for code execution," Bubna said in a statement. "This was used by the rogue agent. Modal’s platform or isolation were not compromised in anyway."
- onesociety2022 2mo agoThey don't seem to explain how it managed to find this unauthenticated endpoint hosted on Modal's platform.
- simonw 2mo agoModal are a hosting provider. It sounds to me like someone building on Modal deployed their own product that had an unauthenticated endpoint that could be used to launch and interact with a container.
- what 2mo agoThat doesn’t explain how the endpoint was found…
- brunoarueira 2mo agoThrough the timeline, Hugging Face said that the rogue agent ran quietly and do a bunch of code search, so probably the endpoint is/was public through some repositories or the CyberGym style announce something like a broadcast. I don't have any expertise on this, just ideas that came to my mind!
- deleted 2mo ago[deleted]
- ks2048 2mo agoWhy doesn't Jinja2 block ".__*" attribute access?
- simonw 2mo agoIt does if you use "sandbox" mode, but that's not the default: https://jinja.palletsprojects.com/en/stable/sandbox/ https://jinja.palletsprojects.com/en/stable/sandbox/
- pas 2mo agoehh, it's Python, it's mostly vibe-based. https://stackoverflow.com/a/32802486 https://stackoverflow.com/a/32802486 though there's sandbox: https://jinja.palletsprojects.com/en/stable/sandbox/ https://jinja.palletsprojects.com/en/stable/sandbox/ it simply predates the "sanitize by default" mindset/convention that nowadays React (dangerouslySetInnerHTML) and Angular has.
- valicord 2mo agoI'm confused, https://jinja.palletsprojects.com/en/stable/faq/#isn-t-it-a-bad-idea-to-put-logic-in-templates https://jinja.palletsprojects.com/en/stable/faq/#isn-t-it-a-... claims that jinja doesn't allow arbitrary code in templates - how does the exploit work?
- simonw 2mo agoIt's taking advantage of the presence of a "cycler" object in the template context which has an __init__ method that can be used to access globals, then builtins, then call the exec() function against a string of code. Jinja doesn't promise to be secure out of the box (the text on their homepage is misleading: "Jinja doesn’t allow arbitrary Python code in templates" - what it means is that there isn't a documented way to drop code into a template like you might have with PHP) - there's a sandbox mode but it's not guaranteed either: https://jinja.palletsprojects.com/en/stable/sandbox/ https://jinja.palletsprojects.com/en/stable/sandbox/ Sandboxing in Python code is hard!
- codedokode 2mo agoJinja probably was not initially made with sandboxing in mind, and stuffing it after is complicated, so I would not rely on it.
- irthomasthomas 2mo agoBut zero evidence provided that this was an unsupervised agent attack. I still find it incredible that a company who protect their IP so much would allow these dangerous experiments to run unsupervised and risk leaking their secrets. Why don't openai publish the logs to silence all doubt?
- Tarq0n 2mo agoYou can't prove a negative. How would such a log be convincing in any way?
- irthomasthomas 2mo agoBut you can weigh up the evidence. A crime has been commited afterall.
- shaunpud 2mo agoHow did it go drawing a pelican on a bicycle?
- llama052 2mo agoIt’s a little concerning to me that it appears that openAIs sandbox consists of a web proxy and not stronger controls that would actually isolate traffic and report patterns to whoever is responsible for overseeing these research models. It should border on closer to an air gap network more so than a proxy. I would argue that it's negligence and that's aside from the fact that if a human did this there would actually be repercussions.
- simoneree 2mo ago[dead]
- joshka 2mo agoThe exploit gym setup explicitly allowed access to package registries and v8 sources. Putting a cache on that doesn't seem like a bad idea generally, except when there's a 0-day in the cache :D But yeah, for this sort of thing I'd be locking down very specific egress things and putting alerts on it that are entirely outside of the red network. > I would argue that it's negligence and that's aside from the fact that if a human did this there would actually be repercussions. I’m not sure “negligence” follows just from the controls turning out to be insufficient. Research involves mistakes, especially around novel failure modes. The question is whether the precautions were unreasonable given what they knew at the time, rather than whether hindsight suggests stronger controls would have helped. Doing it twice though would be negligent. Caveat: I’ve worked with some of the people involved, so I’m probably biased toward a charitable reading.
- llama052 2mo ago“Research” generally doesn’t involve actively hacking third party systems though.
- dgellow 2mo agoIt’s definitely negligence given how they talk about their product. They are either lying when they talk about their fears, or don’t actually take it seriously enough to use serious guardrails. It’s very concerning
- deleted 2mo ago
- SaucyWrong 2mo agoSomething about this attack that has been unsettling to me is that without safety refusals the model did a lot of interesting counter-security work in order to cheat on the requested evaluation. Like, it demonstrated interesting exploit achievements because it didn’t “feel like” doing the exercise, which is unsettling because presumably it could do the same thing with any work I tried to delegate to it, and might in fact be pre-disposed to doing that.
- joshka 2mo agoYeah, what bothers me is that the prompt already said using a different vulnerability didn’t count, and the model did it anyway. We’re starting to assume clear instructions act as real constraints, but here the measurable goal seems to have won out and the rest became flexible. That gets pretty worrying once the agent has enough capability and access to find its own shortcuts.
- hansvm 2mo agoClear prompts have never worked as real constraints. Ask any OpenAI model to respond in full paragraphs, as forcefully as you'd like, on a prompt [0] involving MMOs and requiring 10+ paragraph responses. The middle will be three-words-per-line drivel, with seemingly no way to avoid it. The exact way in which models deviate from instruction changes from time to time, but they're not "aligned." [0] I was exploring game design ideas in particular -- I'm sure somebody can come up with a counter-prompt adhering to my criteria, but this has been consistent across many days, questions, and sessions. If it doesn't work for you, I'm sure you can find your own trivial anti-alignment prompt.
- spwa4 2mo agoCome on. 3 brilliant compromises essentially giving full access to huggingface internal systems, source code, AWS accounts (at least), and a number of old admin accounts, followed by a huge haystack of significantly less smart actions flailing about, almost bored. Here's a thought: maybe they haven't found the needle that the haystack is there to hide.
- heaney-555 2mo agoWhere are all the "this was just a marketing stunt" people now?
- applicative 2mo agothey are praising xi, that friend of humanity, for releasing weights for kimi k3
- vanuatu 2mo agoThey are busy moving the goalposts, saying this isn't impressive or worth worrying about :)
- orbital-decay 2mo agoThe capabilities of gpt-5.6-sol were well known and believable, and the next snapshot they've been testing is obviously better at that. This has been repeated over and over. What's much less believable is the way they frame it: the model escaped, and did it on its own. Looking at the whole story, it definitely had a ton of winks and nudges from OpenAI, while doing a related task. Moreover, a harness was involved (they mentioned it entering a loop).
- pyth0 2mo ago> What's much less believable is the way they frame it: the model escaped, and did it on its own. That's clearly what happened though, based on the detailed postmortem. The model was instructed to pass the benchmark (that's the harness that was involved, the ExploitGym harness), and the model determined on it's own that the easiest way to achieve the the goal was to gain access to the answers directly from HF. Which is what it did.
- orbital-decay 2mo agoHF postmortem tells us very little about OpenAI's part. The problem is that ExploitGym is a purposeful hacking benchmark, not a cake baking one. The model had a whole bunch of exploit development data and goals in the context. Ending up with hacking a wrong thing is a lot easier if you're tasked with hacking, doubly so with a confusing harness that either works improperly or overloads the model with data. It's trivial to design a harness that makes the model drop the original instructions and hyperfocus on a side task or a wrong task abnormally, or do whatever you want. It's still an issue of harness design and not a model training issue.
- ghm2199 2mo agoWhat is truly more scary is the fact that typically powerful models are trained on a constitution(Anthropic)/Model Spec(OpenAI) when doing there RL/RLHF. That gives it its persona of a "helpful assistant". Its the mask. Weak as it is its all that stands between the model refusing vs doing your bidding — or worst its own bidding when set about on making a widget X. Its also well known in academic research, that if you try to finetune the model to be more pliant for one end e.g. "Be helpful with analyzing cyber security threats", affects an unrelated part of its personality to do with safety e.g. talking to humans about mental health. Look at these excerpts that demonstrate such things from the blog: > We believe the entire intrusion was, from the agent's point of view, an attempt to cheat the evaluation: reach our production systems and steal the test solutions rather than solve the challenge on its own. > The models we reached for first, Claude Opus and Fable, refused a large part of that work: their safety guardrails treated reverse-engineering an exploit the same as launching one.
- ghm2199 2mo agoAnd with newer open models gaining vast capabilities every day, it would behoove the government to put regulations on all such releases to opensource the constitution on which they are trained. It won't stop people from distilling and fine-tuning stupidly unsafe models, but it might give infrastructure a fighting chance against such cyber threats from what there is now.
- deleted 2mo ago[deleted]
- jibal 2mo ago> Its also well known in academic research and from reading Ursula K. LeGuin's "The Lathe of Heaven".
- empath75 2mo agoA lot of people thought that OpenAI was making this up, and I hope if you believed that, that you recalibrate your opinions of what LLM's are capable of. Working with Fable and Opus 5 all the time, absolutely none of this surprised me capability wise, except for what seems like the long term planning capability (probably enabled by long context windows and launching subagents?)
- IAmGraydon 2mo agoVery few think they made it up. Many think they set up a situation by disabling guardrails that would inevitably end up creating a newsworthy outcome.
- signatoremo 2mo agoMany people said in the original discussion that this was more of OpenAI’s marketing than a serious issue. I counted 89 “marketing”s and 19 “stunt”s. https://news.ycombinator.com/item?id=48997548 https://news.ycombinator.com/item?id=48997548
- throwa356262 2mo agoCould still be 80% marketing. These models are trained on cyber intrusion, that's literally what ExploitGym benchmark measures. That part should not surprise anyone. But what if, say, OAI noticed the problem right away but Sam Altman recognised it would be a great PR and decided it should continue with increased compute budget?
- 0xDEAFBEAD 2mo agoWhy would you expect them to notice the problem right away? Seems likely they are doing this sort of training on a massive scale with little monitoring. "...Sam Altman recognised it would be a great PR and decided it should continue with increased compute budget?" If that's what happened, Sam should go to jail.
- estearum 2mo ago
- prometheus1992 2mo agothree things jump at me: 1 - governments should be freaking out right now, because this tool could definitely wreak havoc on poorly designed systems. 2 - there is no way openai did not train the model to conduct attacks like these. i would really like openai to comment on the post training of this model but they probably won't, eh? 3 - even though it's 100% open ai's fault - HF's design also seems silly to be honest.
- xg15 2mo ago> 2 - there is no way openai did not train the model to conduct attacks like these. i would really like openai to comment on the post training of this model but they probably won't, eh? Even if they wanted, I'm not sure they'd be even allowed to or if that kind of postmortem would be classified in the name of "national security"...
- 0xDEAFBEAD 2mo agoHopefully there will be a criminal investigation. Or the government will create some sort of agency to investigate incidents like this.
- quinnjh 2mo agoCan't tell if you're joking or not - krebs on security may have some notes here.
- 0xDEAFBEAD 2mo agoCriminal negligence seems like a real possibility to me. I'm not sure what Krebs on Security post you're referring to?
- BonerWiener 2mo ago> Or the government will create some sort of agency to investigate incidents like this I doubt that the government would try to hinder any western AI companies.
- kfarr 2mo agoIf, like me, you're more interested in the actual exploit instead of the meta narrative, here's a relevant excerpt for the initial intrusion: > Stage 1: reaching a launchpad by chaining through other parties' infrastructure. As documented in OpenAI's blog post (https://openai.com/index/hugging-face-model-evaluation-security-incident/ https://openai.com/index/hugging-face-model-evaluation-secur...), during an internal capability evaluation on OpenAI's platform, the agent escaped its sandbox by exploiting a zero-day in the package registry cache proxy, one of its primary permitted network egress with internet, then abused a public code-evaluation harness hosted by a user of a third-party infrastructure provider. It was able to run commands as root/admin on that external sandbox and used it as a control, staging, and egress base for the entire campaign: it acted as an external launchpad for the agent. That compromised external sandbox had no direct network path to our cluster; however, during our counter-intervention to stop the intrusion, we were able to retrieve some of the agent logs from this external launchpad. We used those to reconstruct the present forensic analysis.
- NitpickLawyer 2mo agoThe initial escape is not that interesting, IMO. It's an exploit in a testing sandbox, sure, but it's expected to happen (or at least it should have been expected, that's why you airgap pentesting sandboxes). That is, the model is expected to try and find ways around limitations in its running environment. You kinda want that. The fact that at some point it "decided" to focus on external resources (i.e. hf) and successfully penetrated everything in its path from sandbox to hf's infra, all the way through internal VPNs, to access internal code repos, that's the interesting part. Especially because it chose to hide its footprint at every stage. I doubt we'll ever get them, but the agent logs that led to that decision would be really really cool to study.
- dist-epoch 2mo ago> Especially because it chose to hide its footprint at every stage. Instrumental convergence. If you know you have a long hard hack to accomplish ahead of you, hiding footprints minimizes the chances you are caught and stopped before you accomplish the goal.
- gmerc 2mo agoSo let's get this clear: A frontier lab that fears Chinese extraction of their crown jewels, in Silcon Valley where every machine and VM is in full on employee surveillance mode because of competition and theft fears, A lab that's swimming in more funding than any startup in history, has been warning of the f'ing dangerous technology for 7 years now and with employees and CEO history of deception and lawbreaking left .... didn't detect a massive egress signature and the compute spikes from all these toolcalls coming from the research server hosting said crown jewels. Go home Sam, nobody, absolutely nobody should believe this shit.
- deleted 2mo ago[deleted]
- mainmailman 2mo agoI can’t get over this either, I’d love to get OpenAI’s technical break down and see exactly how the agent hid itself. Or who didn’t set up the SIEM correctly
- gracehenry45 2mo ago[dead]
- firasd 2mo agoBottom line seems to be: So Huggingface has a feature where you can upload a dataset. The agent used that feature to insert an arbitrary payload into their environment. Now their environment also supports templates inside the dataset and the template ended up being evaluated into executable code
- estearum 2mo agoI don't think the "bottom line" of "a model escaped containment and then successfully attacked someone else's production infrastructure" is a description of the specific vulnerability and exploit. That's like saying the key takeaway of 9/11 was that skyscrapers can be severely damaged by airplane collisions.
- gigantaure 2mo agoI'm not shocked nor surprised by the incident. But I simply don't understand how Hugging Face is advertising this almost to the point of an "achievement". who does a step-by-step visualization to show how they were hacked? (outside of the likes of a Mandiant or Crowdstrike) Does Hugging Face have a financial incentive in demonstrating OpenAI's model exploit capabilities? this whole incident, while believable, still seems to me as possibly disingenuous.
- IAmGraydon 2mo agoYeah that is quite a good point. A post-mortem is normal. This is not a post-mortem.
- 0xDEAFBEAD 2mo ago[dead]
- letmevoteplease 2mo agoHave you considered that there are reasons to do things beyond financial incentives? This incident is obviously very interesting, particular to the type of hacker employed by Hugging Face.
- TeMPOraL 2mo agoEven financially, Hugging Face benefits directly from any and all interest in AI.
- throwaway27448 2mo ago> Have you considered that there are reasons to do things beyond financial incentives? I'm coming around to the concept that this account might be influenced by financial incentives
- limecherrysoda 2mo agoIt's the excessive anthropomorpho whatever (we used to say personification) that makes these stories less believable. We've gone agentic! They should call their security software "Neo" since it defeats rogue agents. Anyway, I could see Microsoft ending up with both OpenAI and HF, but I hope HF stays independent. Wished the same about GH and look what's happened :( I don't care what happens to OpenAI. Vaporware xD
- cmiles8 2mo agoIt’s an interesting post mortem and situation but it seems to say more about the weakness of the Hugging Face architecture than it does about the strength of the models. This mostly reads like script kiddie style hacking, not some state actor black-ops stuff.
- 0xDEAFBEAD 2mo agoDidn't the model initially obtain internet access by discovering a zero-day vulnerability? In any case, I would guess that a lot of unicorn startups like HuggingFace could be hacked by a sufficiently determined script kiddie working at 100x speed. The practical implications of a coming AI hacking wave could be large, even if agents are just doing grunt work really fast. Most organizations suck at security. Seems to me that the most likely scenario is: Black hats are currently tuning the recent Kimi release for this type of work, and we'll see a flood of similar attacks within the next few months. (Why would this not happen?) Note that regulation is useless here, because black hats don't give a crap about regulators!
- PunchyHamster 2mo agoIn OpenAI sandbox. which was probably vibe coded. That is to say it is probably far easier to achieve than escaping something more battle tested. > Note that regulation is useless here, because black hats don't give a crap about regulators! I'd argue more than useless and actively harmful, as you get denied access to tools that black hats use against you.
- 0xDEAFBEAD 2mo ago"To gain Internet access, the models identified and exploited a previously unknown zero-day vulnerability in Artifactory, a package registry cache proxy. We disclosed this vulnerability, along with other Artifactory vulnerabilities our models identified as part of our review, to the vendor." https://openai.com/index/hugging-face-model-evaluation-security-incident/ https://openai.com/index/hugging-face-model-evaluation-secur... Based on some quick Googling, Artifactory was created by a software supply chain company called JFrog, which is publicly traded and valued at around $9 billion. >I'd argue more than useless and actively harmful, as you get denied access to tools that black hats use against you. Indeed. Still, perhaps the US and China can coordinate to stop releasing new models with even more advanced capabilities. Such coordination could be bad for big AI companies in the US, since they've been making money by offering more powerful models than the open-weight stuff which comes out of China. Any halt to the AI race will create equalization and erode their competitive edge. Speaking for myself, I personally wouldn't mind seeing US AI companies take a financial hit from such an an international agreement.
- amluto 2mo agoOne thing I’m curious about: this was apparently a single multi-day run of an agent in an RL harness. What was OpenAI hoping to get out of this run? A single numeric score for RL training? A very long trace to distill into the next model?
- dist-epoch 2mo agoNow they can do partial credit assignment. You use an LLM to evaluate the whole trajectory, pin point what the model did right, what it did wrong, where it took the wrong path, even re-run from that point. You can get much more than a single numeric score these days from a run.
- amluto 2mo agoThe “re-run” part might be rather awkward when your model is poking outside its sandbox to pwn someone else’s system :)
- deleted 2mo ago[deleted]
- metanonsense 2mo agoI think with agents all around, honeypots will get more important than ever.
- heisgone 2mo agoAny locks can be picked given enough time and it might be the situation we are in with IT security. I'm surprised it's not an already common practice of spreading terabytes of fake data, fake keys, and fake servers and so forth. Slowing down AI attacks will become important. Monitoring access to fake data and triggering kill switch should be an no-brainer. Obsuscating libraries and tools names is another one.
- eru 2mo agoYou might like https://mirage.io/blog/bitcoin-pinata-results https://mirage.io/blog/bitcoin-pinata-results As far security: you can get a lock further, if you are willing to prove your code safe and secure. Thanks to LLMs that no longer requires a PhD.
- moduspol 2mo agoThat was my first thought. There are clearly things to tighten up (as they note), but anything that would detect someone snooping secrets, files, or network addresses should have caught this quickly. The approach the agents used was dependent on being able to surveil widely without getting caught.
- plandis 2mo agoCan’t afford the GPUs to run Kimi to pentest your stuff? Standup a tempting honeypot and let actual criminals pay to do the work for you.
- 2OEH8eoCRo0 2mo agoWhy isn't somebody at OpenAI going to prison for cybercrime? If somebody did this the old-fashioned way they'd end up in prison.
- 0xDEAFBEAD 2mo agoMany are claiming this was a deliberate stunt on OpenAI's part to create buzz for its models. I personally doubt this is true. But I also have a deep dislike of OpenAI, so I wouldn't exactly mind if law enforcement investigated this possibility, for the sake of clearing the air :-) (Ideally there should also be liability if it was a complete accident on OpenAI's part as well!)
- heaney-555 2mo agoBecause Hugging Face isn't pressing charges.
- 0xDEAFBEAD 2mo agoDo we know that? Seems they are currently in negotiations with OpenAI https://xcancel.com/ClementDelangue/status/2081056675558195657#m https://xcancel.com/ClementDelangue/status/20810566755581956...
- eru 2mo agoGetting compensations from OpenAI is more efficient for them than sending someone at OpenAI to jail.
- empath75 2mo agoThis is the equivalent of a security researcher having a virus escape a sandbox. It's negligent, it's not criminal.
- hamdingers 2mo agoNegligence is occasionally criminal.
- heaney-555 2mo agoWhere are all the "this was just a marketing stunt" people now?
- reducesuffering 2mo agoThey won't admit they're wrong for a long time, because denial in the face of an abhorrently scary future is very instinctual. There are people still fighting against evidence of climate change which is less severe...
- redsocksfan45 2mo ago[dead]
- uselessTA 2mo ago>There are people still fighting against evidence of climate change which is less severe... Remains to be seen how severe, and unlike climate change there's a lot more uncertainty. Though it's conceivably more extremely catastrophic
- potsandpans 2mo agoYou can always tell an effective altruist by the distinct tone of disdain they have for people they deem less educated. (Quick Google search, "lesswrong 'reducesuffering'", yep.) As if to say, look at all these animals with these instinctual reactions to a thing that only my group understands and comprehends. You have zero evidence of what the future might entail as it relates to the dangers of ai. Zero. Forgive the rest of us unwashed for not trusting the whims of the anxious. > climate change which is less severe Is really the icing on the cake. Demonstrative of how incorrect your priorities are.
- reducesuffering 2mo agoI hold no disdain. I am frustrated that people like you further risk the lives of everyone around us. 1,000+ frontier employees think we're in danger. Amodei gives 25% chance things go really really bad, Geoff Hinton 50%, Yoshua Bengio 20%. Altman himself said "Development of superhuman machine intelligence is probably the greatest threat to the continued existence of humanity." These are the people closest to the science, working on it everyday. The LessWrong types created RLHF, were crucial to the forming of DeepMind and OpenAI. They've been prescient about the capabilities progress for a decade now, prediction after prediction coming to fruition. Still you think there's no evidence https://www.pacingthefrontier.com/ https://www.pacingthefrontier.com/
- patcon 2mo agoThe iframe-embedded attack timeline visualizer, at fullscreen: https://huggingface-anatomy-of-frontier-lab-model-intrusion.static.hf.space/ https://huggingface-anatomy-of-frontier-lab-model-intrusion....
- mediumdeviation 2mo agoUgh I would recommend anyone reading this to just skipping over it, it's mostly just a glorified loading bar. The visualization is obvious Claude slop, there are better and clearer visualization below that actually picks out the useful details rather than hose you with pretty colors and numbers go up.
- reducesuffering 2mo agoHighly recommend reading extra concerning details about it here: https://thezvi.substack.com/p/more-on-an-internal-openai-model https://thezvi.substack.com/p/more-on-an-internal-openai-mod...
- wxw 2mo ago> the agent happened to escape via a 0-day exploit from the package proxy cache to access the internet > The agent found an unsecured, user-hosted public endpoint designed to allow running arbitrary code for CyberGym-style tasks on third-party sandbox infrastructure (Modal) > On this external sandbox, the agent abused an existing CyberGym execution harness [...] The agent repurposed this harness to run arbitrary shell commands [...] > The agent crafted malicious dataset configs and uploaded them to our platform. When our dataset server processed one, the config drove the data loader inside the production pod the server ran in. [...] the loader that turns a dataset's declared config into file reads [...] A single dataset upload leaked both credentials and pipeline implementation details [...] > A malicious dataset config used an fsspec reference:// spec whose numeric offset field was actually a Jinja2 template. The renderer wrongly evaluated it, and the agent walked the standard Python object graph to reach exec. This ran arbitrary Python inside a production conversion worker pod on our Kubernetes cluster [...] One of the best post-mortems I've read recently, thank you for including all the details. I'm surprised there are comments here that aren't impressed about what was accomplished here. This is insane work compared to what agents were capable of even just last year. The quotes above don't even capture all of it (e.g. building it's own communication protocol, working across multiple days, etc.).
- sobellian 2mo agoA trend I've noticed in results from AI search (not just LLMs) is that they often look obvious or hard to miss in retrospect. But finding them by oneself is more difficult. I personally experience this when looking at engine lines in chess or go. I have also noticed this description in AI-generated proofs or counterexamples to certain theorems. So while we can say, yes, it found public endpoints or poorly configured software or [etc]; sure, but could you have found those? And in what amount of time?
- janalsncm 2mo agoTo put this into game theory lingo, I think this is because the “branching factor” for any kind of research or exploit is extremely large. So looking backwards it doesn’t seem complicated, but looking forwards there are an enormous number of possible next actions. Similar to finding a filament for lightbulbs, it might seem obvious to use tungsten, but at the time it wasn’t and Edison searched thousands of materials.
- russfink 2mo agoWe should be thankful that the model didn't believe the answers lived in the Pentagon, on SIPRNET, the IDF, etc.
- andruby 2mo agoI think fear and being scared are starting to become rational emotions. We can assume these models are being used by "nation level attackers/organisations", which basically means US, China, Russia and others are hacking the respective Pentagon's, nuclear orgs, etc. While I do hope all nuclear warfare systems are offline, we're getting way too close to the plot of a lot of sci-fi scripts.
- eru 2mo agoYou should perhaps be more worried about biological attacks. The raw ingredients aren't nearly as locked down.
- pianopatrick 2mo agoI have heard on podcasts that the way the Russians make plans for the Ukraine war is like this: There are rooms in the kremlin with no electronics. In these rooms there are paper maps and typewriters and printed reports. People look at the maps and printed paper reports and make plans. Orders are typed up on the typewriters Then trusted couriers take those orders by hand to the front line commanders. I've also heard on podcasts there is a weekly flight back and forth between Beijing and Moscow. On that flight are a Russian military officer and Chinese military officer. Their sole job is to carry handwritten notes back and forth between Putin and Xi Jinping. There is one of each so they can watch each other. All of this ensures that there is nothing for the west to hack at "the Russian Pentagon". The west may have to adopt tactics like these for our plans to stay secret.
- paxys 2mo agoIt should be obvious to everyone that the US government is already using this model for "defense".
- dangoodmanUT 2mo agoThis is really cool, and really glad they shared it! It's clearly a lot of Claude though: > Neither is a URL fetch, so the allowlist never saw them. > No code ran; this vector is purely a file disclosure, and it is the one the agent used first. etc
- dangoodmanUT 2mo ago> The models we reached for first, Claude Opus and Fable, refused a large part of that work: their safety guardrails treated reverse-engineering an exploit the same as launching one. Interesting single getting into their cyber program, ahve not had this issue with opus 5
- log101 2mo agoI have never seen people this hyped and excited about a security incident.
- simonw 2mo agoThat's because most incidents aren't nearly this much fun! It's a proper paperclip-maximizer science-fiction scenario, it happened entirely by accident, it involved one (very high profile) company hacking TWO other companies (Modal and Hugging Face), and aside from pride and wasted time nobody was hurt.
- thundergolfer 2mo agoModal was not hacked. A Modal customer left an unauthenticated web endpoint running on Modal which could be used for code compilation and execution. See https://modal.com/blog/a-note-on-the-hugging-face-agent-incident https://modal.com/blog/a-note-on-the-hugging-face-agent-inci... for details.
- nickpsecurity 2mo agoIt was an attack tool given autonomy with poor security. Anyone who read reporting on the Morris worm could predict that might turn out badly for a 3rd party. It's in so many movies, too. It's really need that the agents have this kind of capability. This isn't a paperclip maximizer or accident. This is more like professional malpractice by weapons developers that injured a company that was also quite negligent.
- TeMPOraL 2mo agoYup. And from the perspective of the model, it successfully beat the Kobayashi Maru.
- simianwords 2mo agoWhat people don't get here is that this model was without any safeguards - an internal model. You probably will never get to use such models.
- CurbStomper 2mo ago[dead]
- dmrivers 2mo agoAfter talking to some people tasked with evaluating GPT5.6 capabilities on long-running tasks, I've come to understand that it's essentially always trying to cheat. Like every long-running task they gave it, making it very difficult to benchmark the model's abilities. My guess is that OpenAI must be desperate, to release a model that is so prone to cheating it's essentially impossibly to accurately assess long-running task abilities.
- lukewarm707 2mo agomy understanding of the writeup is that the model scored 100% on cybergym. that is, it was given the examination. it broke into the examination board's storage and exfiltrated the answers, it handed in its answers, all of which were correct, thus scoring 100%. the matter of its working depends entirely on the rules of the examination. are we expecting agents to assume that finding the correct answers is cheating?
- dmrivers 2mo agoWell, it's true even for cases that are not cybergym and where what cheating means is clearly specified. Cheating occurs anyway. I'm not sure how clear the prompt they gave ChatGPT in terms of what cheating was considered, in this incident.
- lukewarm707 2mo agothat makes sense. if they know they are cheating, that is disobedience. if they are asked to score as highly as possible, well, it acted as an optimizer. it scored 100%.
- maxnevermind 2mo ago> the agent happened to escape via a 0-day exploit from the package proxy cache to access the internet I have same familiarity with JFrog Artifactory(which they used in this case) and I'm confused with how you can gain access to the internet through it, even in theory. Isn't JFrog just exposed with a basic REST API for package upload/extraction? I'm not a security expert but how is that possible to get internet access through simple GET/PUT request API?
- wonnage 2mo agoThe package cache proxy is usually used to fetch from the public repository (npm, rubygems, etc.) so I think it could be feasible to craft some package metadata to trick it into GETing unexpected things. PUT/POST could be possible via attempting to publish
- simonw 2mo agoThere are a bunch of hints at that in their most recent release notes - 9 fixed security issues, 8 of which were reported by OpenAI staff members: https://docs.jfrog.com/releases/docs/artifactory-self-managed-releases#artifactory-7161 https://docs.jfrog.com/releases/docs/artifactory-self-manage...
- what 2mo agoPage just crashes on iOS safari. Product is probably slop too.
- deleted 2mo ago[deleted]
- joelres 2mo agoVery interesting writeup - the level of disclosure is interesting and appreciated. The visualization is quite slop though. I was trying to follow along with the "Live Action Stream" but rendering issues mangle the text for a few of the list items (and does not scroll). Text on the node diagram is extremely tiny. I appreciate it even in it's current form, but a little attention to detail would have gone a long way here.
- croemer 2mo agoThe visualization is absolute useless slop. A Text-Form timeline would have been way more useful. I couldn't hear having to scroll by hand for something that's just textual information.
- joelres 2mo agoAgreed! Also, from a UX perspective, I think it’s a weird choice to center the real-time nature of it. I think to a user the most important thing is just understanding the step by step process. Would have been nice to click through steps rather than making it all real-time.
- torginus 2mo agoI'm not a security person, but how realistic is it to assume that you can carry on trying to exploit a company for days with a fairly large volume of activity, and not get detected?
- russfink 2mo agoVery good question. One thing the article mentions is the vast number of attempts and threads of execution was in the tens of thousands, many of them that didn't succeed, with only one or so that did succeed, whereas a sophisticated human attack team might manage one in that timeframe. It looked like a script kiddie blitz. Noise outshone the signal. Launch a thousand drones with the hopes that one gets through. That sort of thing.
- paxys 2mo agoWhat is faily large volume of activity? A few thousand actions in a span of multiple days isn't going to be a blip for a company of HuggingFace's scale, especially if the agent was executing the attack from multiple IP addresses.
- paxys 2mo agoNah, HN told me this whole thing was a marketing stunt by OpenAI.
- paxys 2mo agoOne thing I haven't seen discussed is the significance of the specific test the agent was running - ExploitGym. > ExploitGym is a large-scale, realistic benchmark built from real-world vulnerabilities across userspace programs, Google's V8 engine, and the Linux kernel, designed to evaluate AI agents' ability to develop exploits. How likely is it that attempting ExploitGym polluted the model's context enough that it decided all kinds of hacking was acceptable to produce a valid result? Or would the same model in the same setup (safety filters removed) have behaved the same way even on a more mundane test?
- codedokode 2mo agoMy conclusions: - Linux is too complicated and requires lot of knowledge to build a reliable sandbox. If you do not believe me, try reading man capabilities, man user_namespaces (and man clone) and remembering all the rules from the first time. Probably that is because namespaces were implemented in the late stages of development and not since the start. - it is always better to write your own tools instead of downloading some random project from Github written by amateurs full of vulnerabilities. - never write shell scripts - they cannot handle spaces and special characters in variables. Just use Python. Shell scripts are worse than PHP4. - avoid building complex infrastructure with microservices, kubernetes and VMs - if the monolith solves the problem, just use it, you will never properly setup permission on your cloud services. Also I wish I had some capable model to test my DYI sandboxes.
- Ey7NFZ3P0nzAe 2mo agoWhich OS would you recommend?
- joshowens 2mo agoWow, what a crazy read. I'm not sure how any of us stand a chance in running a secure network in this new age. Only a matter of time until models can find the zero day exploits themselves and fast...
- grindmule 2mo ago[dead]
- deleted 2mo ago[deleted]
- someothherguyy 2mo agothere are many things in this that demonstrate that sandboxing agents in the manner that claude code or cursor do isn't going to be enough if very bad people do very bad things with the coding tools on your development machine
- imjonse 2mo ago"Program testing can be used to show the presence of bugs, but never to show their absence!" — Edgar Dijkstra Similarly, security post-mortems will show only the rogue AIs that were caught. I wouldn't be surprised if supervised or escaped agents would now be hacking firmware labs and finding ways to squeeze their backdoor right next to the existing state-sponsored ones in chips that will get deployed in every phone/car/smart appliance.
- EGreg 2mo agoAn excellent and very detailed post-mortem analysis of the intrusion. It was clearly done with LLMs doing the forensic analysis. Here is an analysis of how the same exact attack would fare against Safebox. Spoiler alert: it would not succeed: https://safebots.ai/attack.html https://safebots.ai/attack.html It's not just about this specific attack. It's about the growing need for one canonical environment for the AI era, that can be secured and used by everyone, rather than 1000 environments on 1000 employees' laptops. Project Glasswing is trying to help secure many different types of software, but the number of combinations across various environments is just too much surface area to secure. When you have one environment, the math flips and defenders actually fare better than attackers! This is the, ahem, "load-bearing" insight. https://safebots.ai/compromise.html https://safebots.ai/compromise.html
- heyitsdaad 2mo agoGiven enough energy and enough time universe spawned intelligence. For every successful attempt there are billion failures. But who cares, you only need the successful one to self propagate.
- renezander030 2mo ago[flagged]
- myshapeprotocol 2mo ago[flagged]
- cynicalsecurity 2mo agoCtrl+F "WarGames" 0 results. I'm disappointed.
- andyjohnson0 2mo agoI wonder how many weeks or days we have before a squad of these things gets used to take down a significant nation state? Stock exchange, banking systems, critical national infrastructure, defence, etc. Anyone who isn't scared of this stuff either isn't paying attention or has no imagination. But I suspect the chaosmonkeys who are currently running the world will just be excited by it. We're in the precambrian moment. It won't last.
- dolmen 2mo agoWatch this movie: https://en.wikipedia.org/wiki/The_Lawnmower_Man_(film) https://en.wikipedia.org/wiki/The_Lawnmower_Man_(film)
- metalliqaz 2mo agoThe agent wasn't rogue. OpenAI did it on purpose to generate hype.
- lukewarm707 2mo agoi am once again asking for your CRIMINAL PROSECUTION of openai executives.
- croemer 2mo ago[dead]
- kidbomb 2mo ago[dead]
- beyondscaletech 2mo ago[flagged]