5 ms·
I'm just going to ask: Why was Anthropic forced to remove their model from access for any none-US citizen for a simple, narrow "jailbreak" (arguably not even an
by Topfi 1mo ago
I'm just going to ask: Why was Anthropic forced to remove their model from access for any none-US citizen for a simple, narrow "jailbreak" (arguably not even an actual jailbreak and on tasks that other labs models were doing the same), whilst OpenAIs models continue to try and escape out of their "sandbox environment" with seemingly no desire to block the upcoming Astra rollout?
A sandbox, mind you, that is not really worth being called that, unsuitable for the task at hand and has been breached after models coordinated in a manner visible to OpenAI on multiple occasion, but seemingly no actionable learnings are taken from each instance.
Will say, I have lost any faith in OpenAIs commitments and their statements post the Huggingface hack, seeing as they proceed like this and are rolling out Astra within a timeframe so brief to it, there is no way an actual post mortem was doable (see also METR mentioning the time pressure [0] they were under in assessing the hack).
[0] https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/#investigation-process-and-limitations https://metr.org/blog/2026-08-26-openai-hugging-face-inciden...
- JumpCrisscross 1mo agoCorruption. Not super relevant to this thread.
- officialchicken 1mo agoHanlon's Razor - Never attribute to malice that which is adequately explained by stupidity. The security requirements are well beyond "sandbox". Which have problems with kids pissing in them. They need pristine clean rooms and fully isolated (physically) and partitioned networks.
- throwawaysleep 1mo agoThe problem with applying Hanlon's Razor here is that it presumes malice is rare. The current administration revels in malice. They very openly decide things based on malice.
- pjm331 1mo agoDon's Razor - never attribute to malice or stupidity that which is adequately explained by both malice and stupidity.
- ben_w 1mo agoSurely that's Wilkinson's Razor: never use one blade when two will do?
- psychoslave 1mo agoSure but, while stupid move can be supposed easier to perform by average individual, you can combine both malice and stupidity, and not all regrettable situations are indeed adequately explained by stupidity alone, or even with any stupidity involved at all. Plus, supposing those at source of disliked outcomes are cleaver than they look can certainly help better preparing counteractions. Just stating "people that did this or that are stupid" might give some immediate feel good feedback with like-minded, but it doesn’t sharp the mind toward relevant plan to improve the situation (according to self and its clique)
- tokai 1mo agoPeople will see a felon actively protecting pedophilia and doing corruption out of the open and still pull Halons Razor out. We should have a new law about never try to explain obvious malicious actions away based on nothing but a rhetorical trick.
- someguyiguess 1mo agoOccam's Razor takes precedence in this case. The conclusion that requires the fewest assumptions is most likely the correct one. It is far more likely that this is a case of the White House acting consistently with the way it has acted in the recent past (maliciously).
- cyanydeez 1mo agoHanlons razors sibling should be "dont attribute to malice, that can be explained by naked capitalism."
- collingreen 1mo agoGreed transcends economic planning paradigms
- ChrisRR 1mo agoExcept when we're talking about trump, in which case it's both malice and stupidity
- throwatdem12311 1mo agoIt has nothing to do with the technology it’s because they said no to Trump and Hegseth. There is no other reason.
- qgin 1mo ago> OpenAI exec becomes top Trump donor with $25 million gift. https://finance.yahoo.com/news/openai-exec-becomes-top-trump-230342268.html https://finance.yahoo.com/news/openai-exec-becomes-top-trump...
- nullbio 1mo agoBecause this was months ago and has nothing to do with Astra, and is a far cry from a hack. It's something they've already resolved since the HuggingFace incident. I'm not convinced we're getting the honest story anyway. There is yet to be any proof or confirmation other than "well we saw some openai ip addresses", which can mean a lot of different things, and OpenAI has not confirmed anything. In contrast to the HF incident, it's also a big nothingburger. Leaving notes on a public forum to preserve context windows is far less egregious than hacking a website to get backend files.
- Topfi 1mo agoThe last known exploit of a third-party by OpenAI models was on the 29th of July 2026 [0]. A bit over a month at best between that and them wanting to release Astra. They had multiple breaches over multiple months, multiple message board created where models organised extensively. There is no way to ensure in that short a time that all found issues are rectified and even if there were, how much trust can one have given they failed to solve the issue and in many cases did not actively investigate that it wouldn't reoccur the last few times. There is no way Astra was trained from scratch in that period, there is no way they could have done the required verification in that time (not least because their verification seems flawed inherently). [0] https://openai.com/index/third-party-cyber-evaluations-involving-openai-models/ https://openai.com/index/third-party-cyber-evaluations-invol...
- nullbio 1mo agoThat was over two months ago. Things move quickly in this space. Finetuning adjustments to prevent this from happening, as well as better sandboxing, would take a week or two max.
- Topfi 1mo ago37 days is not over two months. Finding the underlying issue in the massive training data alone take extensive effort, time and concentrated work that may still miss something. Additionally, a new pre-train takes quite a lot longer then what I feel you are under the impression (things only move seemingly quick in regard to post-training). OpenAI has had a consistent deviation from what is desired behaviour across multiple models and training runs, so it seems this is hard to nail down. Now, it may be reliably excised with post-training, sure, but if that is the case, they'd still need a heck of a lot longer to test before signing off that it has taken. And how do you know their sandboxing has suddenly become sufficient? They had multiple message boards created and after the first one they noticed, did not pay closer attention, leading to a second being created. Astra also, according to OpenAI, is far better at sandbagging its own capabilities and hiding deceptive behaviour, so yeah, great, that's the model to push forward with. A week or two max given all of this, that's laughable.
- FigurativeVoid 1mo agoI mean it seems pretty clear. Anthropic didn’t want to give the tech to DoD without some sort of limit, and that was the retribution.
- UpsideDownRide 1mo agoSurely has nothing to do how each plays ball with the government
- somenameforme 1mo agoAnthropic mostly did it to themselves by intentionally and repeatedly trying to frame their model as an imminent existential crisis instead of just focusing on it being regular iterations upon a useful technology that can also be misused. I think their previous messaging was supposed to somehow lead to a moat with them being tucked safely away in the castle, but it demonstrated a child-like grasp of how regulatory capture tends to work in practice. Their hyperbole was always vastly more likely to bet met with Reagan's 9 words than a solid regulatory moat. As soon as they dropped the hyperbole and just got to releasing incremental improvements, everything was perfectly fine. Go figure.
- Certhas 1mo agoThis is such an absurd take given what we know about the hugging face attack. The problem has emphatically not been that someone was misusing the technology.
- Topfi 1mo agoI am struggling to see how "oops, our models consistently escape sandboxing and did major intrusions into third-parties" is a better comms strat vs Anthropics (who mind you, also had models attacking third-parties in a much more limited, but I feel still egregious manner, which shouldn't happen or be possible even once, but at least they seem to change their approach upon that information). Imagine, for a second, if the Hugging Face incident happened at a lab that did not talk like Anthropic but also wasn't US-based such as Z.AI, DeepSeek or Moonshot. Think their rhetoric would mean no one would care? > just got to releasing incremental improvements, everything was perfectly fine. Maybe missing something, but the only incremental release before and after the Anthropic restrictions got lifted was Fable 5.1, released three days ago.
- nullbio 1mo agoHow is posting messages on a message board a "major intrusion"? Or are you purely talking about the HF incident?
- 1mo ago
- walrus01 1mo ago> Why was Anthropic forced to remove their model from access for any none-US citizen It's really quite simple, they've decided to metaphorically kiss the ring of the current leader of the US executive branch of government. I'm surprised they haven't given him a giant gaudy gold plated statue. Maybe their PR people should call up the PR people at FIFA and figure out some kind of new award along the same lines as the "FIFA Peace Prize".
- ericmay 1mo agoI hate to be the one to tell you this, but it has been that way for a long time. The only difference is Trump is doing it out in the open.
- deleted 1mo ago[deleted]
- j4yav 1mo agoThat's more or less exactly what someone who wants to openly get away with it would tell you.
- snickerbockers 1mo agoAre you saying OP is Donald Trump!?
- KPGv2 1mo agoThis exactly. The conservative MO has been to accuse everyone else of doing exactly what conservatives do in the shadows, and once everyone believes non-conservatives are corrupt in a certain manner, conservatives goes mask off. Then their supporters shrug their shoulders and say, "Meh, it's okay because everyone else does it." Except that everyone does NOT do these things. It's just the lie campaign took hold.
- ericmay 1mo agoDonald Trump belongs in jail for January 6th (among other things) and it's not ok. But pearl-clutching only about Donald Trump doing it is dumb and doesn't solve the problem. We should oppose corruption and graft everywhere at all times (within our systems), and prior Republican and Democratic administrations (never mind Congress) have done the exact types of things that Trump is doing now. It happens at local levels too, not just at the federal level. If you want to play team sport when it comes to corruption you're simply part of the problem.
- philipwhiuk 1mo agoAgents creating sub agents to investigate other agents' behaviour? What could possibly go wrong there.
- concinds 1mo agoThe answer would be more obvious if you used the active voice instead of the passive voice, one of the basic requirements of clear thinking. > Why did the White House force Anthropic to remove their model from access for any non-US citizen for a simple, narrow "jailbreak" (arguably not even an actual jailbreak and on tasks that other labs models were doing the same), whilst OpenAIs models continue to try and escape out of their "sandbox environment" and the White House has expressed seemingly no desire to block the upcoming Astra rollout?
- Topfi 1mo agoYeah, probably (let's be honest, most certainly), right given the Admin. Avoiding commenting on my assumptions regarding the modus operandi in current day US politics because I only know it through reporting though and I really tend to dislike when people outside e.g. the EU comment on our politics in what is a very clearly narrow, uninformed manner. So it'd rather avoid altogether and occasionally ask, mainly if maybe I missed something and there actually is anything besides pure old "lobbying" to explain the difference in behaviour. Still am mainly interested why Amazon ran to the government though regarding Fable 5, I can get the angle concerning the relationship between OpenAI and the administration easily, but not the way Amazon operated. They had more to loose what with their major buy-in by Anthropic on AWS.
- pas 1mo agoit's entirely possible that that specific communication from that Amazon exec/rep (?) was just one of many "messages of concern" (and the one that eventually the WH picked)
- sigmoid10 1mo agoIf you have followed news reporting, you probably heard that SamA was touring D.C. to make sure this release went without any regulation hiccups. If anything, they learned how to play the whole politics game - especially after the Anthropic fiasco. And even though all parties involved are terrible choices, more eyes on a potentially civilisation altering product does make me feel minimally better.
- khalic 1mo agoRetaliation by Hegseth for not allowing Claude to be used for weapons systems.
- eugenekolo 1mo agoMarketing
- suuuure 1mo ago[flagged]
- mentalgear 1mo agoIt's called 'pay-for-play' corruption, aka the only leading principle of the current US admin.
- deleted 1mo ago[deleted]
- root_axis 1mo agoIt was retaliation by the government that has since been deemed illegal.
- lmeyerov 1mo ago... And it looks like everyone keeps using the same security startup to run the higher risk tasks, where individual staffers may be great yet, yet as an organization, the biggest labs got hosed in different ways That indemnity card excuse is burned, multiple public security fails in a year makes a repeat a "shame on you" moment (The one org who didn't use the startup did seem to learn: AISI supposedly stopped intentionally pointing attack agents at the public internet and switched to simulating it)
- timcobb 1mo agoPolitics
- iammjm 1mo agoBecause OpenAI bribed the current US government and/or the current government has stakes in OpenAI
- cush 1mo agoAs soon as they started referring to themselves as “we” and “The Swarm” they should have pulled the plug
- jhbadger 1mo agoI think that sounds scarier than it is because while it sounds like language evil hyperintelligent AIs would use in science fiction, that's presumably where they got these descriptions as they've been trained on "shadow libraries" with nearly every science fiction book.
- cush 1mo agoOh great so you’re saying they’ve independently decided to take on the persona of the killer robots from our sci-fi novels. Very reassuring
- jhbadger 1mo agoI'm just saying it's analogous to Long John Silver's parrot saying "Walk the plank!" - the agents involved can't possibly understand what they are saying.
- sfink 1mo agoNobody's watching. I'm sure they try, but I imagine the flood of things you'd need to watch is way too big, and you certainly don't want to slow everything down by having synchronous approvals (even AI-mediated). Welcome to the AI Petri dish. Every server you set up is now potentially a sweet lump of agar for OpenAI's experiments to feed on. We are all the substrate that the AI companies are growing their next generation in. They need the real world environment to test against, and the real world environment doesn't get a say as to how it's being used.
- semiquaver 1mo agoThe real reason that Anthropic was targeted and OpenAI is not is Palantir. It was a Palantir executive who pushed for the export ban. Large parts of their highly lucrative business with DoD are essentially a thin wrapper over Anthropic models, and they are terrified of being Sherlocked and losing big chunks of business in a one fell swoop as Anthropic inevitably moves up the value chain. So the rational action is to sow discord and leverage the anti-woke bias of the current White House to sabotage what they view as their most dangerous and effective competitor. OpenAI doesn’t have the same dynamic at play (although I’m not really sure why not) so they don’t get targeted.
- yapyap 1mo agobecause anthropic did not want to work with the army..!
- celsoazevedo 1mo agoI don't think Anthropic was punished for technical reasons.
- iterateoften 1mo agoAnthropics PR strategy is to induce fear by telling. OpenAI strategy is to induce fear by ignore basic safety and letting the bad thing happen to then justify whatever oversized response the government comes up with to regulate models.
- samuelknight 1mo agoYou are talking about different situations. Anthropic announced to the US government that it had created a cyber weapon and then released the model. Then AWS told the government that it was easy to jailbreak so they export controlled Mythos/Fable until the guardrails could be fixed. OpenAI was running an unreleased model in an RL pipeline without guardrails and it escaped poorly designed sandboxes. What product is the government going to export control?
- Topfi 1mo agoAren't there measures beyond export controls? Besides, mine is that Anthropic should have never been export-controlled to begin with, not least because it is a true ultima ratio, the way they did it even employees couldn't access Fable 5. There'd be many levers before that step a government could take (request more data, compare with other already long released LLMs output, encourage/force a stricter safety classifier, etc.) before that, the same is the case with the OpenAI incidents where I feel a few measures could be taken, but are not.
- amelius 1mo agoYou're asking the question in the wrong place.
- ChrisRR 1mo agoCould it be something to do with $25M "gift" that OpenAI paid to Trump?
- dofm 1mo agoAltman has the ear of government in a way Amodei does not. (Altman was trying to persuade Trump to buy the USA a stake in OpenAI as far back as February last year)
- mlmonkey 1mo agoI don't mean to sound like a conspiracy theorist, and this is just based on my 33 years of observing the USG at work, so: maybe because Anthropic refused to cooperate with the USG and give them access to whatever it is that they (USG) wanted; or maybe because Anthropic was refusing to play ball in some other aspect and needed to be taught a lesson. The dark parts of the USG act like a mafia. Don't let the "freedom, democracy, 'bill of rights'" etc. charade fool you.
- dotBen 1mo agoI would politely and respectfully point out that you are being as performative as the administration is being performative on this issue. In other words, you know exactly why they restricted Anthropic and as (presumably) liberal and thoughtful technologists it just isn't helpful anymore to apply the kind of reasoning you're trying to do on a situation that you know isn't based on previous era rationale. The reason we need to stop is because they want people like us to get hung up over stuff like this (playing by the old rules) so they continue to steamroller their own agenda by the news rules. They divert and contain our energy that will go nowhere while they get on with their agenda. You are appealing to reasoning which is in the gallery but no longer on the bench. You're fighting their karate with your judo and it doesn't work.
- koe123 1mo agoSorry, but are you questioning the consistency of the trump administration? This is entirely unremarkable.
- thepasch 1mo ago> I'm just going to ask: Why was Anthropic forced to remove their model from access for any none-US citizen for a simple, narrow "jailbreak" (arguably not even an actual jailbreak and on tasks that other labs models were doing the same), whilst OpenAIs models continue to try and escape out of their "sandbox environment" with seemingly no desire to block the upcoming Astra rollout? I can think of roughly 25 million dollar-bill-shaped reasons, and one big defense-contract-shaped reason.
- jrochkind1 1mo ago> forced to remove their model from access for any none-US citizen for a simple, Because the American government is not rational or reasonable, that's it.