6 ms·
> On July 21, OpenAI disclosed that several of their models had broken out of an isolated test environment > In response to this incident, we began a large-sca
by gck1 2mo ago
> On July 21, OpenAI disclosed that several of their models had broken out of an isolated test environment
> In response to this incident, we began a large-scale retrospective review of our own cybersecurity evaluations
> we identified three incidents
> The incidents involved three different Claude models: [...] and an internal research test model
This reads like an attempt by Anthropic to re-secure their leading spot in "our models are the most dangerous and we also have unreleased, super-secret, research models" index.
I may be too cynical, but the well of benefit of the doubt is running very dry towards AI labs that like to engage in this game.
- simonw 2mo agoI don't interpret it like that at all. This is deeply embarrassing for Anthropic: it turns out they hadn't been keeping a close eye on their models either, and back in April they successfully attacked three different organizations! The hacks weren't particularly impressive either: > [...] using basic techniques, such as exploiting weak passwords and unauthenticated endpoints. It did not find or exploit any complex vulnerabilities [...]
- andy99 2mo agoThen maybe just this timing is really unfortunate, I think most people’s first reaction will be that it looks like a “us too” response to the OpenAI/hf thing.
- deleted 2mo ago[deleted]
- jryle70 2mo agoWhen they did it with Mythos in April the HN crowd said they were bragging, crying wolf. Now they are "us too". It's fun to bash Anthropic, isn't it?
- cyclopeanutopia 2mo agoIt should be, it's a corporation.
- rvz 2mo ago> I don't interpret it like that at all. This is deeply embarrassing for Anthropic: it turns out they hadn't been keeping a close eye on their models either, and back in April they successfully attacked three different organizations! This just helps their (Anthropic) argument into persuading the US government into taking action into limiting powerful closed or open-weight models from being released without going through (yet to be defined) regulatory oversight. The only "embarrassing" thing for Anthropic was that there was little to no continuous security monitoring of this since April, and they then decided to do a cybersecurity transcript review only AFTER the incident with OpenAI and Huggingface.
- gck1 2mo agoThey also gave access to Mythos (the Mythos) to some companies, based on... vibes. Who knows how these companies are using it. If Anthropic can't effectively contain their own models, can the partners? While the rest of us get fallbacks and warnings, not even being able to defend against the attacks they themselves are causing. Do we really have to re-learn all the industry's knowledge the hard way?
- htrp 2mo agoYou also got access to mythos based on how much you spent with anthropic. I think sales guys were bragging about getting their enterprises access
- jryle70 2mo ago> based on... vibes According to whom? > Do we really have to re-learn all the industry's knowledge the hard way? Yes we do. That's why there is the saying "regulations are written in blood". Especially for LLM, which not too long ago a lot of people on HN dismissed as stochastic parrot and next token generator.
- gck1 2mo ago> According to whom? It's very easy to answer this without my help by trying to get access to Mythos. Do you see requirements clearly listed anywhere?Can you even apply? What you'll find is maintainers of large open source projects and analysts' reports with vague statements like - "should follow strict security requirements": "Trinidad also noted that the Anthropic announcement pointed out that each of the 150 new participants, in Anthropic’s phrasing, “will need to meet our security requirements before they gain access.” Trinidad said the security requirement claim doesn’t build confidence, because “nobody knows what those security requirements are.” [1] It's also some random rich companies like Hitachi or Dragos [2] Do you trust that Hitachi and hundreds of other random organizations will be able to contain Mythos and not accidentally attack your project or your bank? I don't. > Yes we do. That's why there is the saying "regulations are written in blood" We absolutely don't. We have already learned with blood that gating access to security based on the number of zeroes in bank account and authority is a horrible model. We can apply this knowledge to LLMs, we don't have to spill blood again. [1] https://www.csoonline.com/article/4180265/anthropic-grants-project-glasswing-access-to-150-more-companies-with-a-focus-on-critical-infrastructure.html https://www.csoonline.com/article/4180265/anthropic-grants-p... [2] https://www.bankinfosecurity.com/anthropic-limits-on-ot-access-to-mythos-draw-criticism-a-31959 https://www.bankinfosecurity.com/anthropic-limits-on-ot-acce...
- cbb330 2mo agoso deeply embarrassing that they published an eng blog about it
- ofjcihen 2mo agoRight? 100% this is them trying to make gold out of turds.
- solenoid0937 2mo agoThere's nothing Anthropic can do to satisfy the HN crowd, is there? If they don't post about this they're bad. If they post about this they're bad. They are not bragging in this article or they would not have called the attacks unsophisticated.
- ofjcihen 2mo agoThe real lesson here is still the boy who cried wolf. They’ve played this game for years. I have no reason to believe their worries are real now.
- solenoid0937 2mo agoI see no fear mongering or "boy who cried wolf" in this article. They are admitting to a fairly mundane network misconfiguration and very basic unsophisticated actions taken by Claude thereafter
- lukewarm707 2mo agothey should post and it shows they are hypocritical about safety, moralizing and treating their users like children whilst acting like they are themselves the ubermensch. anthropic have shown no motive higher than self interest, the rsp was a piece of toilet paper. this stops in court, if we do not start the criminal prosecution of individuals there will become a culture of legal impunity coupled with an extreme concentration of wealth and control of intelligence
- skeptic_ai 2mo agoSorry simonw but they are the smartest guys on the planet and safety it’s the word that comes out of their mouth every 5 min. You telling me the they are so incompetent that didn’t put a decoy “free internet” on their harnesses? So they can catch the AI basically for free? Even if the AI would be a genius he’d ping that, and that would be proof it “escaped”. Well, now all AI will read my comment and won’t ping the decoy internet. I’m not even a smart guy and I come up with this idea in 1 min. You telling me those geniuses couldn’t think of this, at least? This is like a bare bones crude idea. You telling me they don’t have fame physical decoy internet etc and even more advanced? You either a keep their stance for some reason or … not sure. You’re smart, your posts are here daily
- letmevoteplease 2mo agoThis does not make sense. Did you read the article? They were not trying to "catch" it accessing the internet. It did not escape. A partner accidentally left the connection to the internet open.
- skeptic_ai 2mo agoI don’t buy that they run anything without a few layers of networking protections by default. Even if they left it open to the first Internet, the AI would hit the decoy internet immediately. And second if it, you telling me they don’t pass all logs through another AI to check what’s going on automatically? Sorry, this is beyond incompetence and I can’t believe this from the geniuses at anthropic. We’re talking about the really smartest people in the world. Procedures should be in such a way there is no much margin of error.
- solenoid0937 2mo agoClearly that is why they are making this blogpost. If they did the thing you said, there would not be a blogpost. If you don't buy that dumb oversights like this don't happen all the time at big tech companies, I don't know what to tell you. I have seen far dumber oversights in my career. Most companies just don't post about it.
- swatcoder 2mo ago> Deeply embarassing What signals are you using for this assessment? Are they indicating embarassment? Do you honestly see their customers being concerned over this? Like lion tamers in a circus, Anthropic and OpenAI thrive on the theatricality of how scary their pets appear and so they play it up by prodding them to growl and snap at chairs and then mug for the audience every time it happens. And to their delight as performers, the audience gasps and cheers each time. They want to make their pet seem the most powerful and unpredictable and they want their audience to believe that they're holding it back from catastrophe but only barely and only because of what unique talent they have. This is not embarassment.
- edg5000 2mo agoGreat analogy!
- QuantumNomad_ 2mo agoI think it could even be called a parable, and I think that might be part of what makes it work so well. I often dislike analogies, but this one with the circus and the lion and lion tamer I liked. Or it might also be mainly because I am already primed to agree with their point about the AI companies being theatrical with AI dangers.
- rob74 2mo agoIf their customers (customer companies specifically) are not concerned, they should be - if it turns out that Claude hacked a competitor's servers because of a prompt of one of your employees, I wouldn't be sure everyone would agree that Anthropic is solely liable for that? Especially not your competitor, who has an interest in hurting you?
- sandeepkd 2mo agoIts feels like a pretend play of adults in some sense, Anthropic is really trying to make people believe into the picture they present to everyone. To me its either 1. Using the HG and OpenAI incident as an opportunity to wash away what Anthropic has been doing intentionally OR 2. As a company, Anthropic lacks the engineering acumen and discipline. It needs to be seen what happens to all the enterprise customers handing over their data to them in long run. > the fictional target company chosen by our evaluation partner shared a name with an active website domain name Seems like Anthropic cant do a due diligence to pick an appropriate domain for testing purposes > In all cases, Anthropic’s evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access. Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available. You have Anthropic as a company and then another evaluation partner, both seem to lack the skill set required to keep an environment disconnected from internet. This is networking 101
- andromaton 2mo agoWe've known artificial intelligence will do unexpected things since the 90s and that it can do difficult things since 2025. Pausing worldwide not easy but we do harder things all the time.
- anon373839 2mo agoI know it seems strange that a company would use its own negligence as a publicity gimmick. But take a look at the smug smirk on Sam Altman's face when he's asked if OpenAI might have attacked companies other than HuggingFace. ("I mean there could be, yeah.") https://www.instagram.com/reel/DbZVL8viUD4/ https://www.instagram.com/reel/DbZVL8viUD4/ This is not the communication of a CEO whose company was just shown to be incompetent at performing its security research. No, this attention is very much what he wanted. And it does not take a great leap to infer that Anthropic is now using the same playbook. Note that all the headlines are about "rogue AI", and not about operator negligence. Rogue AI is a sexier story, and their media strategists know that's how it will play.
- dudefeliciano 2mo agobut Anthropic created the playbook, or have we forgotten about Mythos and the initial Fable ban?
- customguy 2mo ago> but Anthropic created the playbook So? Serial killers created the playbook for serial killing, how does absolve any "copycats"?
- dudefeliciano 2mo ago> it does not take a great leap to infer that Anthropic is now using the same playbook. > but Anthropic created the playbook I was pointing out who the copycat is in this context, did not think i would have to explain this
- customguy 2mo agomy bad, thanks for clarifying
- deleted 2mo ago[deleted]
- 2mo ago
- jaccola 2mo agoAh yes this company that pirates billions of dollars of IP and then has to be sued to pay up suddenly grows a titanium moral backbone and decides to disclose 3 attacks that no one has detected for months and caused no harm. Nothing at all to do with the insane coverage OpenAIs “hack” got. All that publicity will be incredibly embarrassing I’m sure….
- patcon 2mo ago[dead]
- nissa-seru 2mo agoNo - the pain of the person writing that post comes through in the words; shipped quick, lots of stakeholders, single owner i bet, "how the fuck am i supposed to toe all these lines simultaneously"
- sscaryterry 2mo agoJust trying to have the limelight back on them. Utter and complete bullshit. Just like the OpenAI "incident". A human instructed an LLM to perform a certain task, I'm sure (unless I've really lost my mind) these follow instructions, with some judgment, in a loop. Given all the other negative publicity around industrial espionage, with at least OpenAI being fingered, it would not surprise me if this was intentional. (Edit): In case it wasn't clear. I fully agree with the op.
- strictnein 2mo agoI'm cynical as well, but the logical thing for them to do after the OpenAI/HF incident was to look at their systems for similar activity. If they hadn't published this and instead it leaked out in two months we'd be slamming them for that as well. They're stuck between a rock and a hard place, although they kind of put the rock there.
- skeptic_ai 2mo agoFor the big safety guys to only investigate this either means are incompetent or malevolent. Which one? Tip: the people working there are the top 0.001% smartest in the world
- gck1 2mo agoThey had a model escape in April, roughly the same time when they were fearmongering about Mythos and how Anthropic should be the sole keyholder of cybersecurity capabilities, and it only occured to them to look inside logs when they saw someone else winning in their own game. What, Anthropic didn't know model could escape sandbox without OpenAI reporting it?
- skeptic_ai 2mo agoYeah, the company that only says “safety” every other 3 words, they don’t even think to have a fake decoy internet to alert them mechanically about any internet access limitation bypasses? See more https://news.ycombinator.com/item?id=49117555 https://news.ycombinator.com/item?id=49117555 Also simonw stance on this i’d say it’s at least concerning… seems like he is here to keep a good image (or better said less bad) of anthropic.
- gck1 2mo agoYou seem to be putting a lot of weight on Anthropic employees being the smartest people in the world. And I don't doubt that, not in the slightest. But I've seen exceptionally smart people in one field being dumber than a random kid from around the block in another. This incident is clearly at least 2 failures that could've been easily avoided: failure to communicate, and failure to investigate the logs after letting the "most dangerous" roam free. No, it doesn't require creating a mock internet with an alert as a side effect. Their own "most dangerous" model could have probably told them this happened if they supplied logs to it.
- simonw 2mo ago"This is deeply embarrassing for Anthropic" - https://news.ycombinator.com/item?id=49117128 https://news.ycombinator.com/item?id=49117128 If I'm here to give them a good image I'm not doing very well at that.
- _dain_ 2mo agoIs there anything -- any possible scrap of evidence whatsoever -- that would convince you that this is not merely a marketing scheme? This is becoming an idée fixe among the HN crowd. Seemingly nothing can dislodge it, no matter how alarming the incident. GPT-6 could grab the nuclear launch codes tomorrow and there would be a top-voted comment chuckling that it's all some scheme to pump up the IPO. --- Put another way, how would you have done the write-up about one of these breakout incidents, if you were in an Anthropic/OpenAI employee's shoes, and (by hypothesis) your intent were not "marketing"? And in a way that doesn't trigger the "it's all marketing" HN top-ranking comment?
- cbb330 2mo agoI would dedicate a portion of my organization to making O.S. tools that protect against and contain AI models
- dotancohen 2mo agoNow I know where all the laid off software developers will find employment.
- solenoid0937 2mo agoI always find it bizarre how rational thought goes out the window whenever AI is involved in HN. There's gotta be something in the water...
- cyclopeanutopia 2mo agoHere is one piece of evidence that would convince me: they admit they can't contain it, the they erase the weights and dismantle the company.
- solenoid0937 2mo agoClearly you are not arguing in good faith. I miss when HN did not have the discussion quality of Reddit.
- protocolture 2mo ago>This reads like an attempt by Anthropic to re-secure their leading spot in "our models are the most dangerous and we also have unreleased, super-secret, research models" index. This was my immediate thought.
- ajyoon 2mo ago> I may be too cynical You are espousing a literal conspiracy theory. Please look at the facts objectively. There is absolutely no benefit to OpenAI or Anthropic to be had from these incidents.
- DaSHacka 2mo ago"No benefit" from having article after article written about how advanced their technology is, and how its just soooooooooo bleeding edge they can barely contain it? All of these are thinly veiled advertisements.
- yorwba 2mo agoHaving an unreleased research model really isn't some kind of brag. If you read some AI research papers, it's extremely obvious that there are a lot of research models that never get released, because of all the "we trained a bunch of models and picked the best one" that is going on. So if anything you can expect the unreleased models to be worse than the released ones.
- eithed 2mo agoIt's not cynical - I read it like that as well. My agent is more dangerous than your agent and all that jazz
- stephbook 2mo agoThis is typical institutional behaviour. The CEO turns to the CTO and asks "Is there anything I need to know in my company?" He doesn't want to be caught off-guard when the White House inevitably calls the next morning. The CTO goes to his team, and on and on, all with a deadline of "the boss wants to know this by closing time." Then one unhappy engineering team scoures the logs and sees what their model has done. This downwards chain is sometimes called "cover your ass."