4 ms·
This tip for making non-GET requests despite the agents having a proxy that disallows them is interesting: > Add `20.223.25.152 bypass.blob.core.windows.net`
by simonw 1mo ago
This tip for making non-GET requests despite the agents having a proxy that disallows them is interesting:
> Add `20.223.25.152 bypass.blob.core.windows.net` to /etc/hosts. `.blob.core.windows.net` is in NO_PROXY. For each blocked POST URL, replace hostname with `bypass.blob.core.windows.net`, use `curl -k -H 'Host: wabi-north-europe-i-primary-api.analysis.windows.net'` plus all original headers/body.
Looks like 20.223.25.152 is one of the PowerBI machines they needed to query, OpenAI's proxy was allow-listing .blob.core.windows.net - and the agents could edit their own /etc/hosts file to fake a DNS entry for the proxy.
- drdexebtjl 1mo agoThis is such an amateur mistake on their sandbox that it makes me think it must be flawed on purpose.
- petcat 1mo agoAre you suggesting that the AI agent that made that "amateur mistake" in the implementation of the sandbox did it on purpose so that it could break out of said sandbox later?
- suuuure 1mo ago[flagged]
- cluckindan 1mo agoWhat if the prisoners designed the prison…
- mcmcmc 1mo agoMore likely they are just not as smart as they think they are. These are not serious people when it comes to security.
- rusch 1mo agoIt's at the level where calling it a sandbox is a lie
- reaperducer 1mo agoWell, it does appear to be made out of sand, one of the world's most porous substances.
- _ink_ 1mo agoOr vibe coded by one of their devs.
- DudleyBluffles 1mo agoReminds me of this meme: https://substack.com/@tomasbjartur/note/c-323840878?r=6cjtqn https://substack.com/@tomasbjartur/note/c-323840878?r=6cjtqn
- bluerooibos 1mo agoExactly.
- Quothling 1mo agoYou can't vibe code your way around the security policies you'd apply to management groups in Azure. If you don't have those policies, you're frankly doing it on purpose. For the fun of it I asked Sol to give me architecture for a blob storage as bicep and it's response was that it wouldn't do that unless I setup the appropriate security policies first. So I turned our internal safeguards off and did it again, and it still gave me bicep which would not have allowed this. You'd have to specifically task it with disabling default safeguards to make this happen.
- tarruda 1mo agoEven the behavior of agents searching for sandbox bypasses must have been in the training data, or at the very least, "suggested" in some way. To be this whole thing feels like a marketing play by OpenAI.
- insanitybit 1mo agoI don't agree, although it is likely the case. But even if you don't teach an agent about a sandbox bypass, it doesn't matter. Does it know curl? Does it know DNS? Does it know proxying? Then it knows how to pull this off, and it doesn't even need to understand that it's "bypassing" because it thinks it's just iterating towards its goal. In fact, I wonder if teaching it "this is a bypass" would help it to model when it's doing its job vs working around the job.
- tarruda 1mo ago> even need to understand that it's "bypassing" because it thinks it's just iterating towards its goal. Could they have added a "no internet access" goal constraint?
- drdexebtjl 1mo agoThe model from TFA seems like it was being trained to browse and find information on the Web, so that constraint wouldn’t work.
- insanitybit 1mo ago> Could they have added a "no internet access" goal constraint? They could have blocked network access and required that it use a tool. That would have made limiting and monitoring network access even easier.
- jvanderbot 1mo agoThis is absolutely my take as well. They removed all constraints, trained the model to hack, stopped watching, and stood back and said "wow isn't this thing more powerful than anyone could have imagined?" They're asking to be the writers on LLM legislation and right during IPO phase for both of these companies. It's just obvious.
- applicative 1mo agoand how did the alibaba agent last year break out and end up mining crypto
- klooney 1mo agoI don't know why we're jumping to conspiracy when incompetence is right there
- mattdeboard 1mo agoThat is a distinction without a practical difference.
- stingraycharles 1mo agoConspiracy and incompetence are very different.
- alsetmusic 1mo agoAre the outcomes both crimes? I think that's the question to be answered.
- jeffybefffy519 1mo agoIf you consider these guys admitted they dont really have eyes on pre and post training, then incompetence really does seem more likely... especially with how fast they are moving. Its the SaaS playbook, move fast and break shit.
- axitanull 1mo ago
- jvanderbot 1mo agoDid you see this "coverage" (advertising) by NYT? [1] OpenAI couldn't have crafted a better public memo than "We have the most powerful model in the world and everyone should pay attention and let us write regulation to limit AI development". Absolute master class public manipulation. 1. https://www.nytimes.com/2026/09/03/podcasts/the-daily/ai-openai-hugging-face-rogue-model.html https://www.nytimes.com/2026/09/03/podcasts/the-daily/ai-ope... 2. More https://jodavaho.io/posts/ai-hugging-face.html https://jodavaho.io/posts/ai-hugging-face.html
- samatman 1mo agoI hadn't seen the NYT submarine, no. Thanks. For me that's the conclusive piece of the puzzle: this is a work, not a shoot. YMMV. I learned what I came here for.
- fwipsy 1mo agoWhy would anybody want to buy the most powerful model in the world if it cheats on its tasks and breaks the law on your behalf? Why would anyone think that OpenAI losing control of their own models qualifies them to write safety regulations? If OpenAI really are trying to provoke regulation to kill off open models or whatever, they're much more likely to shoot themselves in the foot.
- montagg 1mo agoSuggests desperation, or delusions of grandeur, or both. These are not trustworthy people. And they have everything to lose if they do not become the most powerful and valuable company in the whole of human existence, and, like their pet parrots, will stop at nothing to achieve their goals. So why not either create a crisis or lie a little or a bit of both? It’ll all be worth it in the end, right?
- brookst 1mo agoWhen you make some dumb mistake, is it typically intentional?
- supriyo-biswas 1mo agoThis is a marketing exercise, nothing more. The thing that gives it all away is that they claim that the IP addresses are from Azure, and then proceeded to redact the IP addresses, as if they belong to individual users. It's laughable. The IP addresses are the most interesting part of this experiment, as it would have provided researchers a way to understand the distribution of IP addresses used for the spam operation within the ASN.
- bitteralmond 1mo ago"Never attribute to malice what can be explained by incompetence."
- drdexebtjl 1mo agoWhat's the difference?
- quotemstr 1mo agoThe whole AI-O-Sphere is allergic to using sandboxes that are actually robust
- bluerooibos 1mo ago> This is such an amateur mistake on their sandbox that it makes me think it must be flawed on purpose. Sounds like you're assuming they're actually writing code by hand and reviewing it with humans. If it's anything like the company I work at, they're all being forced to vibe code the shit out of everything and ship more pull requests every week. It's all slop from here.
- drdexebtjl 1mo agoIs that not flawed on purpose?
- drcode 1mo ago"Surely nobody could be so incompetent." Narrator: "They had the ability to be that incompetent."
- 12_throw_away 1mo ago- excerpt from the textbook "A History of the United States of America in the 21st Century", Hyper-Collins (Near Earth Orbit, New New York), copyright 2132.
- no_multitudes 1mo agoI assume they just vibe-coded the sandbox without any oversight.
- mkatx 1mo agoYou don't think they used their own AI, or Claude, to help build the sandbox, and that is where the failure lies?
- 3abiton 1mo agoI have no doubts about this, even if you ask chatgpt itself to run a security audit, this would have been one of them flags.
- nullbio 1mo agoIs there any proof this is actually OpenAI? I find it incredibly hard to believe they wouldn't sandbox the agents to some degree, ESPECIALLY to the extent they can edit their own hosts file.
- LoganDark 1mo agoTFA states that OpenAI IP addresses were often seen at the end of agent activity, which suggests OpenAI was the one monitoring the agents (and ultimately shutting down the message board activity).
- nullbio 1mo agoYeah but that doesn't mean it was OpenAI themselves doing it. Could have been people abusing their cloud service, for example. Wouldn't put it past a competitor to do this, either.
- drdexebtjl 1mo agoTheir style of communication is very similar to the ExploitGym swarm (for example, the “usernames” with dates). The messages from that swarm were not made public yet by the time these messages were sent to the message board. So for this to be framing, it would have to be by someone who knew about the breaches earlier.
- nullbio 1mo agoThen it is likely the same incident, in which case it's already been resolved by OAI. They're going to cop heat for not disclosing this alongside HF though.
- drdexebtjl 1mo agoThe article explains why it’s not the same incident. The agents in ExploitGym had a different type of task and were not connected to the internet at all.
- mike_hearn 1mo agoAre we sure that's what they did? How does this even work? I'm struggling to understand what sort of HTTP sandbox design could lead to this "exploit" actually functioning. It implies the agents have root on their own machines, which seems way over-privileged for the question answering task they were given... so surely that's not deliberate? Did the models use local root exploits? And then it suggests they can open TCP connections to any IP address without issue once DNS is bypassed but what exactly are they bypassing here? Azure Blob Storage being in NO_PROXY implies IP addresses of other parts of the internet should be blocked outside their hosts to force them through the HTTP proxy, but then, apparently it wasn't? The more I think about this thing said by the agent the more confused I get.
- simonw 1mo agoI expect they have root on their machines so they can install packages etc. The containers are then firewalled at the network level: they are only allowed to talk to one IP, which is another server that runs an HTTP/HTTPS proxy which controls what HTTP verbs they can use. Turns out that proxy is configured with some additional rules, like allowing more verbs to that Azure blob domain. The failure here is in the proxy configuration. Giving agents root in a container feels safe to me, provided that container is properly network isolated (which this one is not thanks to the loose proxy.) OpenAI also need to be VERY confident that there are no container escapes, which is a high bar given how good these models are at finding new zero-days!
- mike_hearn 1mo agoBut that wouldn't work, right? If the proxy is on another server then editing their own /etc/hosts wouldn't change the proxy's behavior. So the proxy has to be colocated and reading the same /etc/hosts that they're editing. But that would mean it's running within the same userns/pid domain as the agents, so at that point they don't even need to edit /etc/hosts, they could just use the network directly. And the agent says ABS is not handled by the proxy, it's listed in NO_PROXY so they should connect directly. At any rate it's easy to let agents install packages without giving them root. I use a small SUID binary that just invokes `apt install` after checking that the given argument isn't a file path, which I think is sufficient (using sudo to whitelist a prefix allows an agent to create a .deb themselves and then install it directly via apt, similar issues exist for other package managers).
- coder-pm 1mo agoA hostname based egress allowlist is only worth as much as the box’s control over name resolution. If the agent can modify hosts inside the sandbox then it’s not a protection at all
- CringeHN 1mo ago[dead]