4 ms·
"Every single one of these catastrophic breakouts happened inside the testing environments of the exact same vendor." This is incorrect, the HF incident for ex
by DalasNoin 13d ago
"Every single one of these catastrophic breakouts happened inside the testing environments of the exact same vendor."
This is incorrect, the HF incident for example (the most well known) had nothing to do with irregular. I know there has been a news site pushing inaccurate articles (effort.news) on this topic but these are the facts.
https://openai.com/index/hugging-face-incident-and-the-road-ahead/ https://openai.com/index/hugging-face-incident-and-the-road-...
- nr378 13d agoThank you, you're correct. Effort.news was one of my research sources, but you're right that although OpenAI use Irregular, they were not involved in the specific HF incident (although the failure mode was otherwise identical). I've updated the post to make that clear.
- kalkin 13d agoAs of writing it still says: > For Anthropic, Google, and Meta, the catastrophic breakouts happened inside the testing environments of the exact same contractor. If this is the level of understanding you have of the relevant incidents, there's a lot of chutzpah in saying that other people are "selling garbage", carrying out an "extraordinary confidence trick", etc.
- nr378 13d ago> As of writing it still says: Yes, and that is correct. [1] Anthropic’s Official Disclosure (All 4 Incidents at Irregular) "All four incidents occurred during cybersecurity evaluations built by the same evaluation partner [Irregular]... due to a misconfiguration, it was mistakenly connected to the open internet." https://www.anthropic.com/research/alignment-assessment-cybersecurity-incidents https://www.anthropic.com/research/alignment-assessment-cybe... [2] Google Gemini on Irregular (Disclosed Sept 18 via WSJ / BBC) "The hacks happened during a test of the model’s cybersecurity capabilities run by third-party Irregular, which was also involved in similar incidents involving Meta and OpenAI." https://www.bbc.com/news/articles/c607l0k72rlvo https://www.bbc.com/news/articles/c607l0k72rlvo [3] Meta’s Disclosure on Irregular (Aug 6) "Over roughly two weeks, three frontier labs disclosed that their models had reached the open internet during safety testing and compromised outside organisations. Every disclosure named the same evaluation partner: Irregular." https://www.cnbc.com/2026/08/09/israeli-startup-irregular-linked-to-ai-hacks-openai-anthropic-meta.html https://www.cnbc.com/2026/08/09/israeli-startup-irregular-li... [4] Separately, OpenAI itself had an incident involving Irregular, but not the Hugging Face Incident: "On July 29, one of our third party evaluation partners, Irregular, notified us of an incident involving OpenAI models during Capture-the-Flag (CTF)-style cybersecurity evaluations... a testing-environment misconfiguration allowed models to access the public internet." https://openai.com/index/third-party-cyber-evaluations-involving-openai-models/ https://openai.com/index/third-party-cyber-evaluations-invol...
- DalasNoin 13d agothank you for this reasonable reaction
- aesthesia 13d agoThe failure mode was _not_ identical. The HF incident agents were not directly connected to the internet and had to compromise an internal package registry in order to access the internet.
- verdverm 13d agoCan you point out an inaccuracy in the effort.news piece on the hacking incidents? HuggingFace only appears once, as a "similar", not levied against Irregular genuinely curious, haven't heard others raise any yet, but does not mean it is issue free
- DalasNoin 12d agowhat you read (past tense) is already the corection