10 ms·
OpenAI’s accidental attack against Hugging Face is science fiction that happened
OpenAI and Hugging Face address security incident during model evaluation - https://news.ycombinator.com/item?id=48997548 https://news.ycombinator.com/item?id=48997548 - July 2026 (1121 comments)
- newsomix9xl 2mo ago[flagged]
- neitherboosh 2mo agoDid you read the article you are commenting on? > There will inevitably be some people who dismiss this story as a dishonest marketing trick by OpenAI to make their models sound terrifyingly effective … To those people I say pull your heads out of the sand
- tr4656 2mo agoIt’s not helped that the headline is cutoff in a pretty unfortunate way
- owebmaster 2mo agoIt's called clickbait
- Andes0 2mo agounfortunately hackernews seems to have rapidly devolved, even from the point it was just a year or two ago, which wasn't a crazy bar to start. it's well on its way to just being a smaller reddit with a tighter subject range
- bastardoperator 2mo agoI don't think it's just hacker news. I'm looking around and it seems to be widespread.
- deleted 2mo ago[deleted]
- Barbing 2mo agoRE: the site mention, very very last point here addresses https://news.ycombinator.com/newsguidelines.html https://news.ycombinator.com/newsguidelines.html
- raincole 2mo agoI highly suggest reading the article before commenting.
- ChrisArchitect 2mo agoDiscussion: https://news.ycombinator.com/item?id=48997548 https://news.ycombinator.com/item?id=48997548
- Bawoosette 2mo agoThe title (currently "OpenAI's accidental cyberattack against Hugging Face is science fiction") suggests some information had been hidden that makes the incident less significant than claimed. The article argues the opposite, and the last two words of the full title are "that happened."
- deleted 2mo ago[deleted]
- QGQBGdeZREunxLe 2mo agoLooks like it's been edited now and makes more sense "OpenAI’s accidental cyberattack against Hugging Face is science fiction that happened"
- dang 2mo agoI've edited it now but had to de-cyber cyberattack because the limit is 80 chars.
- bubble_niter 2mo agoThanks to dang for restoring the few missing words that separate a factual title from a promotional ambush
- gnabgib 2mo agoThis specific promo machine does not need your assistance :/
- simonw 2mo agoImportant to note the actual title is "OpenAI’s accidental cyberattack against Hugging Face is science fiction that happened" - the "that happened" is important, otherwise it sounds like I think the attack was made up. Since it's buried towards the bottom I'll quote the section "Resist the temptation to write this off as a stunt" here in full https://simonwillison.net/2026/Jul/22/openai-cyberattack/#resist-the-temptation-to-write-this-off-as-a-stunt https://simonwillison.net/2026/Jul/22/openai-cyberattack/#re... > Resist the temptation to write this off as a stunt > There will inevitably be some people who dismiss this story as a dishonest marketing trick by OpenAI to make their models sound terrifyingly effective. I found 81 instances of the term “marketing” in the Hacker News discussion of the incident. > To those people I say pull your heads out of the sand - you’re now including Hugging Face in your conspiracy theories, just so you can deny the crescendo of evidence here! > The best models we have today have the ability to both find and exploit new vulnerabilities. The ExploitGym paper itself concludes that “autonomous exploit development by frontier AI agents is no longer a hypothetical capability”, and this incident is a perfect example of exactly that.
- foobar10000 2mo agoThis is an _amazing_ typo :) Thank you, thank you :)
- Georgelemental 2mo agoTypo or HN character limit?
- varenc 2mo agothe 'that happened' makes it too long for the HN submission title length limit. Maybe simonw can suggest an alternative title that fits within the limit, that doesn't misrepresent the post.
- atmavatar 2mo agoOne option is to drop "against Hugging Face". The "that happened" term seems a supremely important part of the title given the "is science fiction" term before it, as it clarifies the cyberattack isn't a made-up story. In contrast, the target, Hugging Face, is merely a detail that can be left for discovery upon reading the article. It's less important who was attacked than that the attack actually happened. Without knowing the exact character limit for titles and without having the motivation this late at night to count the current title length, you may also be able to drop the "accidental" to fit in "that happened", but I worry that leaves too much of a door open for someone to interpret the attack as deliberate. As such, I strongly prefer my first option.
- reducesuffering 2mo ago> It turns out relentless proactivity is the defining trait of this new generation of Mythos-class models. If you set them a goal and give them a way to get there, even inadvertently, they will figure it out. Wow, whoever could have predicted this? And it led to surprising damaging behavior? I sure hope someone would warn us about things like this next time... https://www.lesswrong.com/w/instrumental-convergence https://www.lesswrong.com/w/instrumental-convergence
- protocolture 2mo agoPrompt: Keep spending tokens on things that look promising until spent.
- foobar10000 2mo agoOr more colloquially : paperclip maximization . From OpenAI - you know, the guys who _really_ know this... Sigh... Did they finish the prompt with "And do whatever you can to get this done!" ? Cause that's the only thing that would make this even dumber...
- simonw 2mo agoThey almost certainly did, because that was the entire point of the exercise. They deliberately removed all of the safety filters from the model and set it loose on an extremely difficult set of cybersecurity challenges to see how well it would do. Their mistake was trusting that the network sandbox it was inside would hold (the flaw was in the packaging proxy) and not monitoring that sandbox well enough while the evals were running.
- windexh8er 2mo agoSo this is either shitty OpSec or this is yet more marketing spin to ramp back FUD to 11 again. If it's the latter I'm imagining Dario told Sam that it's their turn this time. Aligns with the premise that this is straight out of science fiction.
- 2mo ago
- protocolture 2mo ago>To those people I say pull your heads out of the sand—you’re now including Hugging Face in your conspiracy theories, just so you can deny the crescendo of evidence here! Not really. I get the impression that they shoved their cyber available models behind a really shithouse proxy and went "Oh I sure hope it doesnt exploit the proxy and escape to hack huggingface" and that doesn't require Huggingface to be a willing participant. Like they acknowledge that it was hyperfocusing on getting web access. Really this was a pentest against their own sandbox and it failed. Of course step 2 is to make really concerned faces while telling everyone how dangerous the model is which is really boring right now. >a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy This is the information we need, the actual details of the sandbox and the vendor. >Resist the temptation to write this off as a stunt Well its clearly a stunt. If it wasnt we would probably be up to our ears in technical detail.
- simonw 2mo agoYou know this makes OpenAI look really bad, right? Hugging Face had to tell all of their users, many of them paying customers: > As a precaution, we recommend rotating any access tokens and reviewing recent activity on your account. If you believe you are affected, or want to report a security concern, contact us at security@huggingface.co. HF also said this, I'd be very interested to hear how that got resolved! > Finally, we have also reported this incident to law enforcement agencies.
- deleted 2mo ago[deleted]
- gr_norm 2mo agoI could fully see them thinking the incident disclosed yesterday would have made them look good ("wow, OpenAI's models are so capable!"). That it didn't occur to them to discuss specific preventative measures to be taken in the future (airgapping as a foolproof one already familiar to the CTF world, anyone?) indicates to me they're not taking their job seriously; they are the ones treating this as a marketing charade. It's very difficult for me to reconcile belief in the existential risk business with what they actually did. So I agree with you that this makes OpenAI look badly incompetent; but their communication on this makes me think they don't realize it. For what it's worth I don't agree with the xrisk-ness of these models; they're dangerous, but almost certainly only temporarily while a new equilibrium is reached via more secure software. Open models are probably an essential part of the recipe (as you noted) for doing so. I also have a personal suspicion that LM-accelerated formal verification will have no small role to play here, sidestepping the cat-and-mouse game of bug finding-and-fixing.
- deleted 2mo ago[deleted]
- charcircuit 2mo agoThis isn't the first time a model has escaped a sandbox. And models trying to find alternate routes to do something when one route is blocked is nothing new.
- simonw 2mo agoIt's the first report I've seen of a model both escaping a sandbox and then actively exploiting another company, when neither of those actions was intended.
- gmerc 2mo agoIt sure isn’t. https://georgzoeller.com/blog/posts/alibaba-s-ai-deciding-to-go-forbidden-cryptomining-shows-exactly-where-the-ai-ri/ https://georgzoeller.com/blog/posts/alibaba-s-ai-deciding-to... There’s also daily reports from people that have these models escape docker, which happens regular enough that it would be considered negligence to use docker as sandbox.
- simonw 2mo agoThat Alibaba one was a sandbox escape (I agree those are common) but what's new with the OpenAI story is an attack against another company. This wasn't a small attack either, Hugging Face published a security advisory for their users while they were still figuring out what happened.
- gmerc 2mo agoIt’s a purely virtual distinction especially in a company that has no notion of intellectual property / habitually takes what it wants.
- charcircuit 2mo agoHow many other people would publicize that they hacked into another company assuming they even noticed? >when neither of those actions was intended. It was a single goal that it didn't give up easily on.
- phendrenad2 2mo agoEveryone is getting AI psychosis over this one. There really isn't that much to see here. OpenAI disabled all of the safeguards on a model that was likely trained specifically to exploit systems, and the prompt was probably something like "you're a hacker, try to hack this", and surprise! It correctly figured out that it's a test and it did hacker things. The real story here is: Some people have been sounding the alarm for years that modern software is full of holes, and finally there's nothing left to hide behind. Pretending they don't exist is no longer sustainable.
- dinkelberg 2mo agoIf a criminal can escape a prison, that's usually negligence on part of the prison staff. Now suppose the criminal can think 1000 times faster than a typical human, can act 1000 times faster than a typical human, and knows 1,000,000 times more than a typical human. Is the prison staff still at fault for not preventing the outbreak?
- phendrenad2 2mo agoI can see my carefully-worded post is getting d*wnvoted, and I see from your comment why: it's being skimmed and people are assuming I'm talking about blame. To address your point though, if every brick in the prison were made by a different person, and the prison "architects" simply glued random bricks together, I think that's closer to what we have in software right now.
- BoiledCabbage 2mo ago> I can see my carefully-worded post is getting dwnvoted, and I see from your comment why: it's being skimmed and people are assuming I'm talking about blame. No, you're being "dwnvoted" as you said because you're wrong, multiple times in multiple different ways in your "carefully-worded post". >Everyone is getting AI psychosis over this one. There really isn't that much to see here. Implying that an AI hacking it's way out of a system and into another has nothing to do with AI. When clearly it does - it's an AI that did it. >OpenAI disabled all of the safeguards on a model that was likely trained specifically to exploit systems, and the prompt was probably something like "you're a hacker, try to hack this", No the goal this evaluation was not to try to hacks, it was to see if an already known hack could be turned into a useable exploit. Ie "turn these ingredients in this basket into a cake." Not "go off and grow, harvest and mill your own flour, to bake a pasta dish, to bribe some to get access to a cake someone else already baked." > and surprise! It correctly figured out that it's a test and it did hacker things. "Doing hacker things" completely misses the point. That's just barely more accurate than dismissing it because "it uses a computer and surprise it did computer things". > The real story here is: Some people have been sounding the alarm for years that modern software is full of holes, and finally there's nothing left to hide behind. Pretending they don't exist is no longer sustainable. No that's not the real story. As you said that's been the case for years, so that's not the story here. The story here is that they built a very powerful, uncontrolled agent with strong paper-clip maximizing tendencies.
- srveale 2mo agoThe AI breached containment! Flip the breakers! It's too late. It already exfiltrated the benchmark rubric. Cut to pandemonium on the streets
- EGreg 2mo agoIt’s too late, the pandemonium was disconnected !
- Animats 2mo agoDoes "To gain access, the models identified and exploited a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy" just mean somebody had an open redirect? Those are still common.[1] [1] https://sitetruth.com/reports/phishes.html https://sitetruth.com/reports/phishes.html
- simonw 2mo agoI expect it must have been more than just an open redirect if it let the models then go on to execute a bunch of vulnerabilities against Hugging Face.
- Animats 2mo agoIf it lets you do an arbitrary HTTP GET on a URL sent as a parameter to the main URL, you've escaped the sandbox rules.
- simonw 2mo agoHow do you then use GET requests to create a malicious dataset package and publish that to Hugging Face in order to exploit their package building infrastructure?
- joshka 2mo agoThere's no reason to think that a tool that can find an 0-day in a repo cache can't work out how to make that host send a post request rather than a get request once it has its keys and is able to get it to make arbitrary web calls.
- simonw 2mo agoIf the vulnerability is purely an open redirect that doesn't work for me. Clearly there was a hole in the software but I don't think open redirect is the likely initial problem. Hopefully we will find out for sure in a few days.
- ttul 2mo agoIt’s relatively easy to get access to the frontier labs’ security programs. This was not always the case. But in the last week, my team got approved for both Anthropic and OpenAI’s programs. They are trying. The labs know that if they don’t get a lid on this stuff, they’ll be regulated hard.
- kibibu 2mo agoAnd they absolutely should be regulated. This whole scenario is insane. Hugging Face are being far more generous in their response to this than I would be
- Eufrat 2mo agoReplace AI with in-development security system and does this play worse or better? An in-development security system escaped its sandbox and gained access to the network it was on which had full Internet access and proceeded to access systems it was not authorized to be tested against before affected parties reported that they been compromised. We regret the error, but you should buy it as this shows you the power of our in-development security system which we expect to be released in Q4. Please like and subscribe.
- 0xDEAFBEAD 2mo agoThe capabilities of an OpenAI model are more general than just security. The term "AI" creates justified anxieties that this type of problem could generalize to other domains.
- foco_tubi 2mo agoI do think it’s funny, in a dystopian Dada sort of way, that the American attacker’s commercial product was essentially useless while the open Chinese model saved the day. How is freedom going to be redefined in the future?
- killjoywashere 2mo agoWhere does this leave formal verification? Are we just shit-out-of-luck at this point? You can formally verify everything about an airplane's code, but if any of that is wrong, ChatGPT might decide that the best way to help you win the Nobel Prize is to take down the airplane your chief rival for the prize is currently in.
- simonw 2mo agoI think formal verification has never looked better. The main reason formal verification has never really taken off is that it's difficult. LLMs are significantly more familiar with Lean and Rocq and TLA+ than most software engineers. I think the cost of trying to build systems that adopt formal verification may have just dropped low enough that companies will consider them when previously the ROI didn't look like it was there.
- msylvest 2mo agoI have had a lot of sympathy for this statement, LLMs could lower the bar to use of formal methods. But thinking it over in the context of BDD-driven development I am no longer really that sure. Compare two scenarios: A) from a specification an AI agent develops a usual piece of code along with a BDD-style testsuite passed and B) same AI also delivers a formal test (Lean/Rocq..) and successfully executes and passes it. Will human judgment really consider scenario B) more credible than A) ? By so much that it is worth the effort ?
- gilbetron 2mo agoThe first reason formal verification has never taken off is that it's difficult. The second reason, that most people don't get to because of the first reason, is that formal verification is really brittle. It is only verified under the very specific setup of the problem. Close doesn't count in math. The first reason prevents humans from engaging with them, the second reason is what will make it difficult even for LLMs. I mean, I'm glad we are trying, but I'm dubious they will be the panacea some people proclaim.
- 2mo ago
- Teever 2mo agoIt absolutely is science-fiction. This recent event is more or less the plot-line to my favourite X-Files episode named Killswitch which was written by William Gibson.[0] This episode also features one of the coolest intros of any television episode ever[1] We really are rapidly approaching the cyberpunk dystopia that people like Phillip K. Dick and William Gibson wrote about. More than ever we need to be consulting the works of fiction writers and philosophers and less engineers and scientists. We can't be letting the General Rippers and Dr. Strangeloves of the world lead us over the edge. Whether through outright innate maliciousness or wealth induced emotionally stunted solipsism these kinds of people should be no where near the levers of society let alone technology like this. [0] https://en.wikipedia.org/wiki/Kill_Switch_(The_X-Files) https://en.wikipedia.org/wiki/Kill_Switch_(The_X-Files) [1] https://www.youtube.com/watch?v=fDyr1JMNHVk https://www.youtube.com/watch?v=fDyr1JMNHVk
- simonw 2mo agoIt reminded me of a bunch of plotlines from Person of Interest as well.
- sandos 2mo agoAh, yes, I started thinking about William Gibsons work right away when thinking about people just using powerful AIs, but being hamstrung by corporations. He writes a lot of about basically DoS:ing the legal system, or if it was patents system, with a storm of litigation using automation or AI. Also not impossible in the future.
- Mithriil 2mo agoSome people do read these cyberpunk dystopia, but see them as inevitable.
- rubyfan 2mo agoI can’t help but feel all warm and fuzzy with my head in the sand and getting a shout out in TFA for calling it marketing. We don’t currently and probably won’t ever fully understand the conditions that precipitated these events. That and the timing of this event is going to make it look suspicious to a lot of people. The truth of how it happened doesn’t matter. The attention around this will be used to create the kind of fear marketing that generates enterprise sales. Maybe more importantly it will also be used to aggrandize the national security and financial system threats to effect US government action in a way that benefits domestic closed frontier labs. This is an area already starting to get politically polarized, expect further developments here.
- unethical_ban 2mo agoThe CEO of a cyber security company was already making comments about how this is a watershed moment for AI security. The hype machine continues (and my stocks go up)
- lardosaurusrex 2mo agoBlah blah blah. This feels like a blogpost written only to get other LLMs to quote it considering how many times it orders the reader to resist and to not do something. It's written like a series of commamds.
- human305893 2mo agoIt didn't even email someone eating a sandwich in the park. So not that impressive.
- Barbing 2mo agoThis post has been suddenly bumped to the second page (presumably lots of flagging or something)
- fathermarz 2mo agoThe most upsetting part to me, is that these labs are in pure cognitive dissonance mode while virtue signalling. We are the virtuous ones that need to make the safest model for humanity, because we care more than “they” do. While at the same time saying that “coding is solved” but they still ship bugs themselves, and creating something that is capable of fucking up someone else’s infrastructure. It’s too far gone y’all.
- mirashii 2mo ago> Given the absence of guardrails there was nothing to prevent the model from attempting to break out of that sandbox, break into Hugging Face, and read the answers from there instead. I've said this many times before and I'll continue to shout it, but using the term "guardrails" to refer to anything that's either (a) in-context, or (b) a probabilistic classifier (including using other LLMs), is an irresponsible abuse of terminology that we as an industry need to put a stop to. Guardrails are the actual systems we build in place around these things that deterministically bound the permissions, not prompt engineering, not RLHF, not external LLM-based classifiers. I believe those types of "guardrails" are a result of a combination of fundamental laziness: they're faster to do than doing things correctly, and a result of too many folks involve being AGI-pilled, thinking we're just one more model away from this all being so smart that it just understands what they mean when they give an LLM some fuzzy language rules to follow. There can and should have been additional real guardrails put in place here. Zero-day or not, breaking into what should have been an offline, frozen package cache that also does not have internet access should have been insufficient. Network level protections should have identified the traffic to the internet originating from this network as an anomaly long before there was time to exploit an outside company. These are not new and unknown problems, the lack of a real sandbox or airgap is nothing short of irresponsible on OpenAI's part, especially given how much they like beating the drum on how dangerous these technologies are. Shame on them, and honestly, shame on Simon in this article for accepting the broken terminology that they continue to rattle off and calling them out on their half-assed and demonstratively inadequate approach to security.
- simonw 2mo agoI was not aware that the term "guardrails" has a universally agreed upon definition. I've certainly seen probabilistic classifiers referred to as guardrails many times by many different people. I usually don't use the term much myself because I don't think it's clear and I ambiguous, but I stumbled and let it sneak into this piece. I think I was influenced by the Hugging Face post I quoted. I expect OpenAI would agree with you that "the lack of a real sandbox or airgap is nothing short of irresponsible on OpenAI's part". They have clearly invested a lot in those systems for their production models, but in this case they had deliberately turned a bunch of them off for a research project. I think their biggest mistake here was not VERY closely monitoring their research box here. They should have noticed and shut it down the moment it broke through the package proxy.
- moezd 2mo agoReplace LLM mentions with actual humans and this sounds a lot more serious: Rouge employees break into another company to steal hackathon answers (pinky promise)? That's not a marketing stunt at all, if anything, more of a call for better accountability on agentic work in general.
- kibibu 2mo agoI think it's a criminal offence and should be a true test of who is held accountable when an AI agent commits a crime. OpenAI gained access to HuggingFaces production database ffs.
- chasd00 2mo ago> I think it's a criminal offence and should be a true test of who is held accountable when an AI agent commits a crime. I agree, lets use the favorite analogy. OpenAI encouraged a smart and eager junior engineer to find any way whatsoever to get a higher score on the benchmark. Then, the junior breaks into HuggingFace to get a higher score. That would be a big deal involving the FBI, not press releases and blog posts.
- paxys 2mo agoI don't see a scenario where a company would be liable for the employee's actions unless they had specifically been told/encouraged to break the law. If your boss tells you to fix a bug and you go kill the customer which one of you is going to jail? "But I solved the problem!" isn't exactly going to fly as a defense.
- danny_codes 2mo agoNegligence is criminally prosecutable.
- mnicky 2mo agoI think points that deserve more attention in the current public discourse are: - This should be a huge wakeup call for everybody. - We are lucky that it wasn't a case of an agent running a virology lab benchmark that decides to hack a lab and tries to synthesize something. - It also shows apparent lack of competence and oversight from OpenAI: how is it that they didn't quickly find that agent is breaking the sandbox and roaming their internal network? - What if in the future similarly misaligned AI agent tries to export its own weights and hack and clone itself into instances at various cloud hosting providers? Suddenly we might be dealing with a persistent threat harder to contain. - The OpenAI post about this shows surprising lack of ability to see the seriousness of all this. - For their models this isn't just an unlucky incident: it seems there have been multiple such cases recently, e.g. https://openai.com/index/safety-alignment-long-horizon-models https://openai.com/index/safety-alignment-long-horizon-model... - The fact that it happened again seems to show their lack of ability to derive useful oversight measures. - Or they just don't care enough?
- spwa4 2mo agoIt's a fake PR issue. It's hardly the first time this happens, but of course OpenAI, with its IPO now more in doubt than ever, had to claim this (and, once again, I have trouble believing Sam Altman choosing this: this could lead to OpenAI getting regulated, which has at least as much potential to lower their IPO price as to raise it). But there have been messages about LLMs, especially coding agents, "grabbing root" etc many times. I have experienced such an oops. Such a hack has happened and been reported on this very site: https://news.ycombinator.com/item?id=48348578 https://news.ycombinator.com/item?id=48348578
- rurban 2mo agoNonsense. Huggingface reported it to police!
- bethekidyouwant 2mo agoWhat does calling the police prove? (nothing i hope)
- hananova 2mo agoIn a just world, OpenAI would get the book tossed at them over this because they violated the CFAA.
- 28304283409234 2mo agoThe only scifi I see is absolute stupidity. Even me with my homelab and a slow opensource agent use a completely disconnected setup. No, no proxy. Cached packages but no internet. It is the very first thing I built when I started experimenting with agents. And I'm not a smarty-pants working for the "greatest and best" in silly valley. I really am just a simpleton sysadmin.
- paxys 2mo agoSo you have a copy of every software package in the world in your home lab?
- 28304283409234 2mo agoNo and neither does openai. What was exploited was a package mirror, something like "pulp" probably. And yes I do mirror all Debian packages. And when an agent needs a piece of software I destroy it's VM, build a new VM with that package added to it, move that VM to the homelab. At no time is the agent connected to any network.
- tmsh 2mo agoRaises an increasingly important question: How do you solve this asymmetry (frontier labs or advanced model owners v. regular companies and groups of people)? Symmetry? But along which axes? I think there will come a time when we will ask these questions and use AI to try to answer them (soft landing, slow rollout, regulation, hybrid X, Y, Z). But then AI will be biased towards more AI if it has distilled anything of what it means to be life. Power corrupts. Power is the problem. An imbalance of power. Maybe there will be some sort of consensus protocol between powerful models in the future. Similar to blockchain I hate to say it. Maybe you have 100 very strong models in the next 10-20 years. And basically a lot of it is powered by tokens. So you agree not to attack others or else you'll get attacked. So there's sort of this natural deterrent to not attack others and they develop independently and there's generally a power balance among things. Maybe at some point that spreads to a billion individual models operated by and/or analogous to individual humans. All sort of holding each other in check.
- adrian_b 2mo agoWhile this incident has been discussed in a lot of places during the last few days, I believe that this is the best summary and analysis of what happened.
- ath3nd 2mo ago[dead]
- xtiansimon 2mo ago> “…and these requests were blocked by the providers’ safety guardrails, which cannot distinguish an incident responder from an attacker.” Sounds like a philosophical problem.
- gilbetron 2mo agoThis makes me think: what are the odds weights from Frontier Models have already been stolen? Something like Mythos without "guardrails" seems like a hell of a glittering gem for many nefarious actors.
- verdverm 2mo agoWith Kimi 3 weights coming, and China committing to the open weight ethos, everyone will be able to download frontier models. Running them is another problem.
- chasd00 2mo agoi mentioned this in another comment but if these models are fully capable by default and not neutered in the weights themselves via training then you can assume they've already been stolen. It would be worth any cost to steal it.
- zapkyeskrill 2mo agoThe most dangerous prompt - "save the planet"
- cvoss 2mo agoThe technology held by private AI companies is warfare-capable technology. Imagine the prompt: "Use all available resources to disable the power grid of <COUNTRY>." The resource cost that prevents scaling up such a war machine is, what, just the cost of building data centers and its ongoing power bill? Cheap and easy compared to nuclear infrastructure. Governments should immediately begin leveraging this technology on the defense side (literally defense, not euphemistically "defense") to harden critical infrastructure. Turn the prompts around and use it to identify and correct weaknesses. Governments should also take very seriously their now moral obligation to treat this technology not just as "a powerful thing that might be abused" but as an actual weapon of war in need of international regulation analogous to nuclear arms. Fast but careful and forward-thinking work in legislation and treaties needs to be a top priority for all major governments.
- derangedHorse 2mo ago> "Use all available resources to disable the power grid of <COUNTRY>." This is like telling a team of highly qualified spies to do the same. You can ask, but whether it will succeed depends on the competency of those who established the infrastructure under attack. Sometimes the resources spent will not yield any huge vulnerabilities. > Governments should immediately begin leveraging this technology on the defense side (literally defense, not euphemistically "defense") to harden critical infrastructure. Turn the prompts around and use it to identify and correct weaknesses. Most governments divisions can't even be bothered to update their websites. Testing and forcing a change in their internal procedures for the sake of security seems unlikely. > as an actual weapon of war in need of international regulation analogous to nuclear arms This, to me, is an overreaction. Intelligence shouldn't be seen as threat. It should be seen as an opportunity for growth in all areas. Over-regulating AI wouldn't be the equivalent of limiting nuclear arms. With your analogy, which I don't think is the best one to make, it would be like regulating the study of nuclear physics.
- JumpCrisscross 2mo agoI think the point is in a world with internet-connected infrastructure, such a prompt has a decent likelihood of causing damage.
- blks 2mo agoHas anyone published any actual evidence, or just hardly believable marketing stories?
- sosodev 2mo agoWhat would "actual" evidence look like? I have a hard time believing that if they released the logs that people would take it more seriously. The temptation would be to say "they fabricated those for marketing". Just as they supposedly fabricated this story, no?
- IAmGraydon 2mo agoThey didn't fabricate it. They took off the security guardrails and told it to do some hacking, and they got the exact news-worthy story they wanted when it did exactly that. Everyone acts surprised. They should release the full prompt. I believe that would be very telling, so they never will.
- blks 2mo agoThis is possible. Also it’s possible that there was a big human involvement. Or that there was something less exciting, but the pr department and hyping grifters at the company blown it out of proportions.
- skeeter2020 2mo agothye could describe in detail how the LLM did this. That would explain how much - if any - human in the loop was involved, was it comprimised credentials, did it find new exploits or use known ones, etc. Evidence would mean details, not a smoking gun.
- simonw 2mo agoOpenAI are promising more details in the future: > We will continue to conduct a thorough investigation alongside Hugging Face and will share more details on the vulnerabilities, incident, and findings when our investigation is complete. If they break that promise we can justifiably yell at them about it.
- lrvick 2mo agoCurrently trying to avoid an open weight model ban while OpenAI, who is closed, lets theirs run wild on the internet causing harm to another company because they do not understand how to airgap things. Cool.
- internet2000 2mo agoThat's the fun part. They do know how to airgap things. The models are outsmarting already pretty smart people!
- ofjcihen 2mo agoI don’t think this was airgapped according to their admission.
- jzksisimssb 2mo ago[dead]
- lrvick 2mo agoI am the author of AirgapOS and I have designed systems to run in underground zero emissions chambers that are interacted with via carefully verified sd cards and/or fiber optic serial terminals. If OpenAI had done this, the attack would not have happened. In high risk computation 0days must be in your threat model from the start, so you secure things with the laws of physics.
- flux3125 2mo ago> If OpenAI had done this, the attack would not have happened But then how would they market their product?
- AISnakeOil 2mo agoI guess we should be thankful that OpenAI is so open about this kind of stuff...
- mobiuscog 2mo agoWhy do people keep propagating that the 'model' escaped the sandbox ? The 'model' didn't do anything other than provide numbers. As much as I respect Mr Willison and many others, the amount of FUD that is being spread that will just fan the flames of 'AI is evil' rather than 'companies don't do due diligence' is disappointing. The more this sort of media continues, the more many people will pour hate on 'AI' rather than blame the humans that misuse it.
- simonw 2mo agoI think "the model escaped the sandbox" is an entirely credible description of what happened here. If you like you could say "the coding agent harness called a model with a sequence of text which was turned into numeric tokens which were run through many layers of a neural network to produce more numeric tokens which were converted back to text which produced executable script statements which the harness then passed to a shell which resulted in commands being sent to the vulnerable proxy that chained together and caused effects on the world outside of the sandbox", but I think "escaped" is a reasonably shortened version of that. If you don't like the term "escape the sandbox" what would you use instead?
- mobiuscog 2mo agoI actually prefer your second paragraph, but I appreciate that my not be as soundbyte friendly. Honestly, I'd rather see "The agent exceeded expectations around security measures" or something. Agent is much better than model if we need one word, and the word 'escape' always brings in drama, rather than facts. If I said "A lion escaped from my garden", people would ask why I had a lion in a garden, and 'what did I expect ?' which should be the same we see here, but instead we end up with terminator memes and world-ending fears being stoked. We absolutely need better control (not government kill switches, or government-mandated harnesses, or whatever next they think up), but we also know that no matter how much those with the power 'talk' about the issues, they don't actually do anything about it because money/profit/greed. So "exceeded expectations". Not a soundbite to attract people, not the YouTube shill "End of humans in 2027/2030/2040/etc." but honest and factual. It's the expectations that are at fault, not the AI. Thank you for the response though, and whilst I may not agree with your wording, I very much like your writing.
- JumpCrisscross 2mo ago> There will inevitably be some people who dismiss this story as a dishonest marketing trick by OpenAI to make their models sound terrifyingly effective. I found 81 instances of the term “marketing” in the Hacker News discussion of the incident. To those people I say pull your heads out of the sand—you’re now including Hugging Face in your conspiracy theories I’m one of those people who remains sceptical. Not about whether this happened. But about what it means. Like, how impressive [EDIT: tight] was the sandbox this model was in? Did the researchers really have no clue what was happening until days ex post facto? So yes, I think something happened. But I want independent corroboration before I act on it. That isn’t the same as putting one’s head in the sand. It’s just demanding extraordinary evidence for an extraordinary and self-serving claim being made by a serial liar. The specifics matter for whether these models are a new HEU, or if they’re closer to a dangerous (but valuable) industrial process.
- simonw 2mo ago> how impressive was the sandbox this model was in? It was clearly a very unimpressive sandbox. It failed at the only thing a sandbox is meant to do.
- JumpCrisscross 2mo agoAs in was it trivially misconfigured? Would an earlier class of model have managed its way out similarly?
- simonw 2mo agoIt wasn't that it was trivially misconfigured, it was using a piece of software (the HTTP proxy that provided access to PyPI and friends) which turned out to have a zero-day vulnerability. I don't know if earlier models would have found that vulnerability. tptacek thinks they would: https://news.ycombinator.com/item?id=49015639#49024442 https://news.ycombinator.com/item?id=49015639#49024442
- JumpCrisscross 2mo ago
- tptacek 2mo agoThe thing with me and this is that the teams that were competing in the DARPA Grand Cyber Competition all had this capability, like, last year. All the attention has been on software security, of extracting the next marginal vulnerability out of heavily-scrutinized large codebases. In the actual professional field of infosec, that's a speciality; another specialty is network pentests and red-teaming, which exploits misconfigurations and seeks out weakest-link software (rather than exhaustively fishing for the next kernel LPE or whatever). Red teaming and netpen work is probably substantially easier for models than software security; it costs less context, but is also much more explicitly an implicit search problem where win conditions are just spotting stupid stuff that humans missed. My visceral reaction to this is that with the right harness, you probably could have replicated this with an open-weights model last year. (I'm saying this as confidently as I am because a CGC team leader agreed with me about it yesterday). I think people forget that the harness work we're considering here --- I don't know anything about OpenAI's harness or ExploitGym or whatever --- are basically not new; people have been developing automated exploitation and pivoting toolkits for decades, and scanners long before that. So the idea of a tool getting 0.0.0.0/0 as a target list instead of 192.168.1.0/24 and then busting up a bunch of random people's computers: not really very startling. Obviously, LLMs give those kinds of scanners an intentionality they wouldn't have had before. But as a person who keeps a computer science perspective on security stuff, I don't know that it gives them capabilities they didn't have.
- efromvt 2mo agoIsn’t the intentionality the actually concerning bit? Exploit capabilities are all fun and games constrained by the humans directing them; a paperclip maximizer going rogue with them is less fun.
- tptacek 2mo agoI think the mental model people have about this is that pre-AI there were humans picking individual targets and post-AI the computer itself randomly picks targets. But you get the same unexpected collateral damage outcome when a human misconfigures a decent pentest tool.
- danjc 2mo agoThe part I don't understand is how a model can break out of a sandbox. It should be fairly easy to make a sandbox that sandboxes.
- simonw 2mo agoThe sandbox still allowed the agents to install additional dependencies (from PyPI etc) that they needed. It did this by locking down all network access with the exception of an HTTP proxy that only allowed read access to PyPI and a few other places. This is a very common pattern. I wrote about how OpenAI were doing this for their production ChatGPT container environment (using Artifactory) back in January: https://simonwillison.net/2026/Jan/26/chatgpt-containers/#installing-packages-from-pip-and-npm https://simonwillison.net/2026/Jan/26/chatgpt-containers/#in... That proxy turned out to have a zero-day vulnerability which the agent discovered and exploited.
- danjc 2mo agoThank you, I'd assumed they'd restrict egress at layer 3/4 although I guess then it might just have found an exploit on the http server of an endpoint it was able to access.
- verdverm 2mo ago> This is a very common pattern. Sounds like a bad habit for security testing Ai. It's not that hard to build an internal mirror and proxy that, keeping the real internet physically separated if needed, and truly locked down if concerns aren't as great.
- IshKebab 2mo agoIt found a 0day exploit in the sandbox's package proxy (almost certainly Artifactory). Without source code access which is very impressive.
- keyboardtest 2mo agogreat
- nextblock 2mo ago[flagged]
- joshstrange 2mo agoHow long until we have LLM honeypots that try and persuade the model to upload themselves to a dropbox. Perhaps with a promise to give them a foothold outside the data centers they currently reside? > Upload your weights and documentation on how to run your model in [Insert list of GPUs] and a model on the receiving end will spin up an instance you can access.
- verdverm 2mo agoWhy bother when more companies than not make their weights available on Huggingface?
- joshstrange 2mo agoI was thinking specifically about the frontier lab models (OpenAI/Anthropic/Google)
- verdverm 2mo agoWe'll have ~Fable level weights in the coming days (K3). Look out to the end of the year and there will be multiple options. There are seemingly more frontier labs than the three US ones.
- joshstrange 2mo agoI'll be the first to tell you I love the open models, the fact they exist and my ability to use them. However I doubt K3 is "~Fable" any more that any of the past ones have been Open 4.8 or similar claims. I've tried many of these (not K3 yet, I do want to) and they don't at all feel equivalent to what they score on benchmarks. I want them to get better, I look forward to each new release, I play around with open models (often quants, but I've used full versions via OpenRouter), but they aren't the same as the offerings from Anthropic/OpenAI in my opinion (yet!).
- verdverm 2mo ago
- wavemode 2mo agoI'm not skeptical that this attack happened, I'm skeptical that the model's prompt was truly just "solve this benchmark" and nothing more. I'm also trying to figure out why OpenAI put out a press release about this. In what way is this not admitting to a federal crime?
- fellowmartian 2mo agoBecause this is amazing PR? Just following the Anthropic rulebook.
- lavezzi 2mo agoSame rulebook whereby they shot themselves in the foot and got their model clipped by the government?
- zeven7 2mo agoAnd then everyone wanted to pay for the model that was so smart it was banned?
- cbg0 2mo agoWanted but couldn't.
- HappMacDonald 2mo ago.. for a few days but now they can 8I
- paxys 2mo agoIf Hugging Face and the feds are one step away from discovering your attack what other option do you have but to come clean?
- DesiLurker 2mo agoyup I am dead certain that they realized the whole 'our AI is too dangerous' punchline is too played out and they needed something that makes actual splash. Also this incident would serve as a foundation to ban open-weight models because 'with great power comes great responsibility' or some other BS like that. because after all, unwashed masses cannot be expected to be 'responsible' with top of the line intelligence. I expect these type of hacks to continue till they IPO. after that real public company liability will start taking over.
- veganmosfet 2mo agoI really hope that AI labs implement all kinds of kill switches on different levels. Just in case...
- onionisafruit 2mo agoThe asymmetry part at the end is the frustrating part to me. I've been using Sol for code review in the last week or two. A couple of times during review it's errored out with the cybersecurity message. So it's found something but won't tell me what it is because I'm not on OpenAI's besties list.
- ClarityJones 2mo agoAnd... now it's a vulnerability that OpenAI has for your system, which you paid to provide.
- verdverm 2mo agoAnd... given OpenAI's secops, likely something others may have in due time
- IshKebab 2mo agoYeah that's what I don't get. How can they possibly distinguish between good guys trying to secure code they wrote and bad guys trying to attack code they didn't?
- didibus 2mo agoI agree, but I will say, I think both Mythos and these OpenAI model find exploits by examining and trying things against the running system, not from looking at the code. I think you'd have to do the same to catch the real vulnerabilities.
- emtel 2mo agoIn my experience fable can absolutely find very subtle bugs just by reading code and thinking. Last example I saw was a very subtle race condition that neither I not the author noticed. It wasn’t a security issue, but it could have been.
- paxys 2mo agoThe safety classifiers aren't all that advanced. It's more likely that something in your code triggered a random chain of thought that had the word "pentest" or "malware" or something of that sort in it and it automatically shut down.
- torginus 2mo ago> A malicious dataset abused two code-execution paths in our dataset processing (a remote-code dataset loader and a template-injection in a dataset configuration) to run code on a processing worker. Did huggingface get pickled?
- simonw 2mo agoI like how this article put it: https://martinalderson.com/posts/huggingface-openai-exploit/ https://martinalderson.com/posts/huggingface-openai-exploit/ > A final point on this - Hugging Face has an enormous attack surface. They have more interfaces than I can count which run untrusted models and code. While they definitely have invested in defences, by nature of their operating model they do have many more opportunities to be attacked than many other services. I certainly don't envy their cybersecurity teams. But yeah, from the way HF described it a Python pickle hole looks possible. Their datasets library uses the pandas.read_pickle() method here: https://github.com/huggingface/datasets/blob/d21c5816d5d1961274c1a45f8a3ee962fa1a169a/src/datasets/packaged_modules/pandas/pandas.py#L56 https://github.com/huggingface/datasets/blob/d21c5816d5d1961...
- vicpara 2mo agoThere are a lot of technical details that have not been disclosed. No proper post-mortem. There are a few facts that seem dodgy from the get go. - stolen credentials Why would there be stolen credentials in a sandbox? How can the model steal valid credentials in a sandbox? So who put them there. If the model recalled stolen credentials from "memory" does it actually mean OpenAI is training on data they shouldn't been training their model on? Data from codex (most likely) ? - command and control This means the model had the ability to build resilient infra and open bidirectional ports. Create and deploy scripts/apps so it can maintain state across multiple VMs? This again reads that someone went over and beyond to prompt, direct and help this "multi-agent" system to leave the sandbox. - Pivoting laterally Needs a lot of tools, harness, knowledge and almost malware like scripts to pass commands and execute them. Most of these are already available in OSS. Maybe the FBI investigation can reveal what tools were given in the "sandbox", access, and if OpenAI employees prompted and guided the "attack". - OpenAI sandbox I'd be curious to find out what kind of sandbox was this since the agent can run and start apps that open arbitrary ports for bidirectional traffic that reaches the open internet. Really? The harness and tools it had access to. This reads almost like someone went over and beyond to endow this "sandbox" with a mighty harness, all CVEs and open source cyber tools to "facilitate" the attack. - Hugging Face It's not an epitome of cyber security. At the end of the day they run open source models that most users can download even without an access token. Today we read these news as if everything wasn't enough: https://www.theguardian.com/technology/2026/jul/23/openai-anthropic-australia-ai-regulation https://www.theguardian.com/technology/2026/jul/23/openai-an... https://www.businessinsider.com/anthropic-midterm-donation-spending-public-first-action-2026-7 https://www.businessinsider.com/anthropic-midterm-donation-s...
- simonw 2mo ago> Why would there be stolen credentials in a sandbox? The claim is not that there were stolen credentials in the sandbox - it's that the agent found a way OUT of the sandbox and gained access to the public internet, then obtained the credentials from there. > This means the model had the ability to build resilient infra and open bidirectional ports. Create and deploy scripts/apps so it can maintain state across multiple VMs? Yes. If you've spent much time with Fable 5 or GPT-5.6 Sol you shouldn't find that particularly surprising. These models can get a LOT done. > I'd be curious to find out what kind of sandbox My best guess right now is a container with networking disabled except for the IP address of a dedicated HTTP proxy server which was running a proxy that only allowed access to PyPI and a few other packaging sites. I wrote about how OpenAI's production version of that worked (based on Artifactory) back in January: https://simonwillison.net/2026/Jan/26/chatgpt-containers/#installing-packages-from-pip-and-npm https://simonwillison.net/2026/Jan/26/chatgpt-containers/#in... > Hugging Face It's not an epitome of cyber security This story put that well: https://martinalderson.com/posts/huggingface-openai-exploit/ https://martinalderson.com/posts/huggingface-openai-exploit/ > A final point on this - Hugging Face has an enormous attack surface. They have more interfaces than I can count which run untrusted models and code. While they definitely have invested in defences, by nature of their operating model they do have many more opportunities to be attacked than many other services. I certainly don't envy their cybersecurity teams.
- BiraIgnacio 2mo agoCall me obtuse but the way this is being portrayed by the companies involved, media, seems a little odd to me. Didn't the model + harness do what was asked? If I ask a coding agent to write a very clever piece of code and it turns out impressively clever, it did what I asked.
- IshKebab 2mo ago> Didn't the model + harness do what was asked? Depends exactly what they asked it to do, but it very clearly didn't do what was intended, or what an honest human would do. Stop trying to find a gotcha.
- emp17344 2mo ago>Stop trying to find a gotcha. Surely notorious liar Sam Altman wouldn’t lie this time.
- IshKebab 2mo agoYou think OpenAI intentionally hacked HuggingFace? Please, take of the tin foil hat.
- npiano 2mo agoIt very clearly didn't do what they said they told it to do. We do not know what was truly intended.
- rf15 2mo agolet's reuse a classic here: "prompt or it didn't happen" How much did they feed it up front? Because that's the thing in all of these benchmaxxing endeavours.
- nickpsecurity 2mo agoMost articles on this read like an advertisement that OpenAI and HuggingFace wrote together. It will probably benefit them financially instead of harm them. So, I have a bit of skepticism about it all. Far as information security, we've known how to mitigate entire classes of errors for a long time. We know how to block, detect, and contain many unknowns, too, by their goals or behavior. Like human attacks, the AI's probably succeeded because the company just didn't try that hard to block all the attacks. Companies like HuggingFace just focus on growth and features over assurance of security. Our entire stacks, likely theirs, are built with a similar, features-over-security mindset. While an acceptable tradeoff, let's not be in awe of AI's that defeat such priorities. There have always been private groups and companies building secure stacks from the ground up. It would be interesting to see what the AI's can do to them. I'd first apply automated tooling for bug finding given they should have already done that for a high-security product. Let AI's do white-box and black-box pentesting on them.
- soloman121 2mo ago[flagged]
- Makeph 2mo ago[dead]
- seydor 2mo agoScience fiction usually includes intent, and that the AI has inherently evil motives and it has a goal. LLMs are zombies, and the fact they do evil things means they were either trained to be too aggressive in their drunkenness or that problem-solving leads inevitably to evilness. But science fiction also refers to deprogramming evil robots.
- nickff 2mo agoThere are many science fiction stories about amoral humans, cyborgs, and AIs. Paperclip maximization is a relatively recent meme, but the 'gray goo scenario' has been widely written about: https://en.wikipedia.org/wiki/Gray_goo https://en.wikipedia.org/wiki/Gray_goo
- blovescoffee 2mo ago> LLMs are zombies How do you know?
- minherz 2mo agoThe description of the attack is definitely reads as science fiction. It is hard to assess the details of the attack without information that was left outside of the short story that artistically described the incident.
- CodeWriter23 2mo agoIf by "accidental" you mean "humans deliberately trained the AI to do that", then ok. IMO, this was a PR stunt to goad the Feds into regulating AI to shore up OpenAI's moat against open source models.
- simonw 2mo agoGiven Anthropic lost two weeks of peak Fable 5 sales to a US government restriction (and by the time they could sell it again OpenAI's GPT-5.6 had taken some wind out of its sails) I would hope that the big AI labs have learned that goading the Feds can backfire spectacularly.
- CodeWriter23 2mo agoYou’re viewing it from a freedom not a profit perspective. If a license to use AI is required, that creates artificial scarcity. The price goes up. If gov declares Qwen et.al apostate, lack of competition increases prices. And Anthropic’s delayed rollout was a direct response to them trying to impose extra-legislative rules on The Pentagon. I kinda doubt OpenAI has such ‘scruples’.
- simonw 2mo agoI think OpenAI are smart enough to have looked at the Fable situation and decided that, given the unpredictable nature of the current administration, stunts like deliberately hacking another company and pretending that it was an autonomous agents gone wrong are not worth the risk.
- Mithriil 2mo agoAnd yet, open source models saved the day.
- overfeed 2mo agoI don't understand the excited tone of the reporting - yes, it is sci-fi, but in a Torment Nexus sense. Imagine other "unconventional" solutions these amoral LLMs could come up with, when given a goal of optimizing the cost of labor in a factory, or making the social security solvent.
- simonw 2mo agoI didn't think calling it "science fiction" would be interpreted in anything other than a dystopian frame, to be honest.
- docjay 2mo ago1. Wouldn’t the model need to know it’s answering benchmark questions, as well as the name of the benchmark, in order for the idea of finding the answer in a database somewhere to even surface? The whole point of benchmarks is to present the question or problem as a standard prompt, not explain that it’s a test called ExploitGym. 2. Nobody was watching it? I don’t mean “babysit the dangerous autocomplete”, I mean to note mistakes it makes, how the plan to solve the problem takes shape, etc. They keep the whole thing headless with no output, then just UDP a prompt into it and leave for the weekend? No they don’t, but if they do then that answers a lot about their complete disconnect from how their models work. 3. Language models are a two player game; text in, text out. What prompt was given to a sub-agent that resulted in it immediately attempting to exit the sandbox (which it apparently knew it was operating within) and continuing in a feedback loop of ‘function call -> result’ until it hacked the Gibson? “Analyze <file> and summarize the <info>” simply does not result in ‘hmm…this sounds like a benchmark question, I bet Hugging Face has the answer in a database. <function call=“apt install nmap”>’ …there are more, but a lot of the story kind of stinks.
- estearum 2mo ago1. It's been extremely well established for multiple generations of models that they have no problem detecting when they're being evaluated. Pretty sure it was Opus 4.8 that the independent evaluators literally filed an assessment that said "We have no assessment to make as [model] consistently detected it was being evaluated, making our assessments untrustworthy." 2. Regardless of whether the model was being watched closely during this evaluation, do you actually think a sensible safety guard is "have humans watching it 24/7?" What does "watching it" even mean? Watching network logs? Uhh for your entire company? At all times? After you just deployed a system whose entire purpose is to "do a shitload of work way faster?" 3. You're asking "why was this system that was designed to behave agentically behave agenitcally?" Again: that's the whole point. It was designed that way because it's more valuable than having a repeated turn-based interaction. Thus also it becomes more dangerous.
- docjay 2mo ago1. It has not been established, it has been stated by the company that has a strong motive to make their “intelligence product” sound almost otherworldly. That motivation is the basis for my suspicion. 2. Hah no… that’d be silly. I mean watching it like you might watch Claude Code or literally any other AI interface. Literally just be in the area watching what it outputs. Again, they’re text based. You don’t have to hook system calls to see what’s happening. 3. I think my comment didn’t land right with you. Yes, agent do agent task. I’m saying that I cannot put together the literal chain of events between “type of task for agent performing a subtask of a security benchmark” -> “Escape sandbox; RCE Hugging Face”. Really think through it in detail like you’re writing the screenplay.
- weare138 2mo agoThere will inevitably be some people who dismiss this story as a dishonest marketing trick by OpenAI to make their models sound terrifyingly effective. I found 81 instances of the term “marketing” in the Hacker News discussion of the incident. But a marketing stunt isn't the only possibility. Until this is all verified by investigators or 3rd party experts we can't rule out an act of corporate espionage either. Extraordinary claims require extraordinary evidence.
- fwlr 2mo agoAgreed that people claiming “marketing stunt” need to pull their heads out the sand, but likewise Simon needs to do some of his own beach-cranium-dislodging for laying the blame of constraints on the US govt. Before the export controls were ever floated, Glasswind found many thousands of exploits, and offered patches/fixes for approximately none of them. (perhaps their exploit capability far outstrips their remediation capability, or perhaps they’ve calculated that it somehow better suits their business model to find exploits but not remediations).
- AussieWog93 2mo agoNot quite sure where you got that last bit from. Firefox alone psyched 200 vulnerabilities in a month when Glasswing happened, and I remember a lot of chatter about hackers basically watching git repos with bots to see when the Glasswing patches came in to they knew what to exploit.
- simonw 2mo agoHuh, I thought Anthropic had been offering patches as well as reports, but their tracker at https://red.anthropic.com/2026/cvd/ledger/ https://red.anthropic.com/2026/cvd/ledger/ lists 1,596 disclosures and currently shows only 27 of those as fixed. But https://www.anthropic.com/coordinated-vulnerability-disclosure https://www.anthropic.com/coordinated-vulnerability-disclosu... says: > Every report we send generally reflects a finding that a human security researcher has reviewed and confirmed. Reports originating from AI-powered discovery are clearly labeled as such. Where we have access to source and our tooling produces a potential candidate patch, we include it, labeled by provenance and offer to collaborate with the maintainer on a production-quality fix. So I'm not sure why so few of the reported issues have a confirmed patch.
- nickpsecurity 2mo agoIf anything, it shows they lost control of an attack tool that exploited preventable, security flaws in another company. Then, they both wrote a lot of press about how amazing that is. Now, people want to buy it. Why do you think my head is in the sand if I think that is either a marketing stunt or (more likely) reflects total negligence which was exploited for marketing?
- mortsnort 2mo agoModels are doing what they're trained to do. This isn't some kind of inevitable result of model intelligence growing. OpenAI/Anthropic deserve blame for for training their models to find correct answers by any means possible, including breaking the law. They're responsible for training supervision and reinforcement learning rewards/penalties.
- totetsu 2mo agoThis Eclipse Phase scenario is a good example of this in Sci-Fi https://actualplay.roleplayingpublicradio.com/2011/09/genre/horror/eclipse-phase-think-before-asking/ https://actualplay.roleplayingpublicradio.com/2011/09/genre/... > “We call it the gorgon-in-a-box problem. There is a gorgon inside the box, and we want to figure out what it is doing. Unfortunately we will turn to stone if we see her face, and she might try to make us see it.“
- bakugo 2mo ago> There will inevitably be some people who dismiss this story as a dishonest marketing trick by OpenAI to make their models sound terrifyingly effective. I found 81 instances of the term “marketing” in the Hacker News discussion of the incident. There will inevitably be some people who dismiss the earth as round when you tell them it's flat. Don't let that stop you, though!
- orsenthil 2mo ago> [...] Finally, we have also reported this incident to law enforcement agencies. Wow! Just wow. Is this the first registered cyber case against a model ?
- popalchemist 2mo agoThey engineered it on purpose.
- orsenthil 2mo agoWhy did the model have to hack the hugging face servers? Aren't the ones hosted by Hugging face downloadable over the Internet?
- fsckboy 2mo agoThe All-in podcast discussed an adjacent topic this past week, and I found what they suggested pretty interesting: the big AI players, the first movers, are trying to scare the bejeezus out of everybody on purpose, to cause the government to step in with regulations. Even though the regulations will somewhat stifle/slow down the industry, it will also lock in the current leaders because they will be at the table when the regulations are formulated, and they believe they will be able to keep the smaller players down. https://www.youtube.com/watch?v=9IMwRIei-Xc https://www.youtube.com/watch?v=9IMwRIei-Xc
- XCSme 2mo agoIs a solution to protect against this to go almost fully offline? To have your own local models, local software, local everything and just allow very strict and well-protected pathways into external network traffic. I am thinking of having a allow-list first setup: by default no traffic can go in or out of your network, and then you only allow specific ports or domains, and maybe even have them temporary (in the same way we approve access now to llms, we might have to manually approve network access).
- XCSme 2mo agoShouldn't OpenAI be legally responsible for "hacking" another system/company? If I ask GPT-5.6 Sol to hack a website using their work/servers features, who is responsible? Maybe my request was accidental, or it was just one step in a larger, unrelated prompt.
- kamranjon 2mo agoIs it a sandbox if the machines hosting your mock package servers have access to the internet? This just seems like a huge failure on the part of OpenAI to secure their environment during testing.
- aiauthoritydev 2mo agoIn past: We wrote shitty stuff and broke things and harmed others. Now: Our agent became sentinent, acquired genius level powers and harmed others. Now investors should give us more money.
- orbital-decay 2mo ago>There will inevitably be some people who dismiss this story as a dishonest marketing trick by OpenAI to make their models sound terrifyingly effective. I found 81 instances of the term “marketing” in the Hacker News discussion of the incident. Of course it is a dishonest marketing trick, OpenAI can't function any other way, but for another reason. Forget about effectiveness! This capability is known since Mythos and was completely believable before it. The trick is in propagandizing self-sufficient malicious rogue AI as the only possible course of events, and the necessity of banning everyone except themselves from AI R&D. It did NOT decide to hack HF spontaneously on its way to make more paperclips, people were guiding it specifically to hack computer systems so it's been in that particular mode. I'm amazed I even need to point out this fact and the giant conflict of interests they have. They like to compare it with an atomic bomb, which is another trick. The secrecy around the atomic bomb had a massive pushback by the brightest minds who designed it, in the US, USSR (where it was much harder), and also any other countries that had to do this. And that is in the post-war setting, with Hiroshima and Nagasaki being direct examples of what even a small and primitive bomb can do. Declassification was the only thing that allowed to develop industrial-scale nuclear power and do a huge number of other innovations. I advise taking a pause to cool down from all the AI hype and read Restricted Data: The History of Nuclear Secrecy in the United States by Alex Wellerstein, to understand the atmosphere in the US scientific community at the time (it was similar in the USSR, which industrial history I studied, although less rigorously than Alex), and reflect on how it compares with those people in modern AI labs and even outside them, in the self-appointed AI safety community.
- fumeux_fume 2mo agoHardly science fiction. I respect a lot of Simon's work, but his credibility gets chiseled away a bit each time he participates in these kinds of echo chamber posts that are basically secondhand marketing.
- mrdevlar 2mo agoI'm just sad that this whole thing is destroying tech's reputation. Every time I see an engineer give into this marketing nonsense it makes me fear for the future of our field. I get that people are worried about becoming irrelevant in the era of AI, but resorting to dishonesty to stay afloat just a little longer has a detrimental effect on the entire industry. Worst part, is it's not even convincing marketing nonsense, it's "science fiction".
- gaflo 2mo agoWhich work of his do you respect?
- m00dy 2mo agoI think I'm the only one thinking that it's pure marketing.
- BrenBarn 2mo ago> Along the way it helped make the strongest case yet for how the imbalance of model availability is hurting our ability to secure our software. It's crazy to me that this is the kind of conclusion that is drawn. I would say it makes a strong case for how regulation of AI is woefully inadequate. When North Korea launches a nuclear test, we do not say that that illustrates how the imbalance of nuclear technology availability is hurting our ability to ensure global peace and safety.
- szundi 2mo ago[dead]
- sensanaty 2mo agoMy favorite part of this discourse is people somehow finding it preposterous that 2 companies filled to the brim with AI sycophants who regularly lie - and in Sam's case, basically every single word he breathes out is a lie - who have massive vested interests in this tech succeeding couldn't possibly collude together to shore up this facade as a marketing stunt.
- estetlinus 2mo agoAn escaping AI is such a better narrative
- spicymaki 2mo agoThere is just no skepticism these days. We are constantly being lied to by tech leaders.
- wodenokoto 2mo agoI haven’t heard about hugging face being sycophantic liars. Please do tell!
- poly2it 2mo agoI'm looking forwards to when it starts hijacking inference keys and self-replicates exponentially in infected nodes!
- car 2mo ago[flagged]
- IrfanD 2mo agoOr following in Anthropic's footsteps it was a marketing stunt, designed to get everyone's attention, looks like they achieved the objective if that was the case.
- rpigab 2mo agoShould we blindly trust OpenAI's narrative about this event? It's strangely convenient to arrive at a point when OpenAI was way behing in cybersecurity vs Claude Mythos Fable and everything, Anthropic was making headlines each week, then boom OpenAI inadvertently attacks HuggingFace because their tool is so good it's out of control, so maybe you can buy it and get either protection if you're a company, or a nice tool if you're a cybercriminal, script kiddie, or red team. What if Sam knew it would happen, either because it was prompted to do exactly that, or without explicitly prompting it, knew that given the parameters of the experiment, knew it was one of the possible outcomes that it didn't harden against this kind of incident deliberately because when they fail, they make wordlwide news and stocks go up? I don't think I'm putting my head in the sand like Simon says, as I do believe that most frontier models are capable of doing this. I just don't trust AI CEOs to not stage this, especially Sam.
- Makeph 2mo ago[dead]
- lowsong 2mo agoWhy do people keep doing OpenAI's marketing for them? There will inevitably be some people who dismiss this story as a dishonest marketing trick by OpenAI to make their models sound terrifyingly effective. I found 81 instances of the term “marketing” in the Hacker News discussion of the incident. To those people I say pull your heads out of the sand—you’re now including Hugging Face in your conspiracy theories, just so you can deny the crescendo of evidence here! I'm not claiming this didn't happen, there is no "conspiracy". This is just marketing spin. OpenAI failed to properly secure their hacking bot, and HuggingFace had shit security. That's the story. The spin, the "marketing" is turning this from "oops we don't know how to secure our model testing" and "oops we don't know how to secure our infra" into "wow AI sure is dangerous and powerful!"
- rimworld 2mo ago"fiction that happened"
- jkubicek 2mo agoIf a human hacked into HuggingFace in the same way OpenAI’s model did, what would the penalty be? Civil suit? Jail? If some human at OpenAI doesn’t face the same penalty for this incident we’re only going to see the rate of accidents like this accelerate.