3 ms·
If they actually wanted to test the model without internet access they'd have run it air gapped, not relied on a buggy software sandbox. This is pretty clearly
by ifwinterco 1mo ago
If they actually wanted to test the model without internet access they'd have run it air gapped, not relied on a buggy software sandbox.
This is pretty clearly a marketing stunt by OpenAI, otherwise the story just doesn't add up
- phatfish 1mo agoMaybe the test/task itself wasn't intended as a marketing stunt. But the response to fallout with "going rouge" certainly was. The joke was the other western "AI labs" had to quickly follow up with their own marketing cover about their "super intelligent" models "going rouge" as well.
- ben_w 1mo ago> not relied on a buggy software sandbox. Third, while we had tested and validated this sandbox, the agents were able to chain together previously unknown vulnerabilities (“0-days”) in the package management service exposed within the sandbox to bypass restrictions How were they supposed to know about "previously unknown vulnerabilities"? > This is pretty clearly a marketing stunt by OpenAI, otherwise the story just doesn't add up The "it's just a marketing stunt" allegations never added up, to me. I've been seeing such claims since GPT-2, where people were laughing at them for daring to practice how to secure a model before it got dangerous, generally by eliding the word "before" in that sentence. Because there's tests other than what the big companies use, we've been able to see for ourselves the rapid improvements at least approximately match what the companies themselves claim with the models they do actually release; and now this unreleased model is able to automate felonies when asked to do so, while the rest of us use the actually-released models to assist in finding bugs and security issues in our own code. Even without that, HuggingFace stated they reported this incident to the FBI before OpenAI knew it was their systems which did it.
- YawningAngel 1mo agoI think a careful and thoughtful person would reasonably expect given the circumstances that the models would look for, and might very well find, vulnerabilities in JFrog to exploit and take action accordingly. For example, it seems as if JFrog itself had broad access to OpenAI infrastructure and the internet: > In the following days, the agents exploited our internal research infrastructure and the Hugging Face platform. On July 9, one agent searched for ExploitGym solutions and stumbled upon an application hosted by a customer on Modal, another AI cloud platform. This application was running “CyberGym,” a related evaluation to ExploitGym. The agent discovered an exploit to achieve control over the workload sandbox and looked around hoping that a previous agent’s evaluation run in the sandbox had solved its ExploitGym task. It did not find anything helpful there, but in the process it established a stronghold in the application from which to launch future attacks. This implies to me that L3 and L7 firewalls were not in place that would have prevented broad access from JFrog. I think a lot of shops would have had those.
- actionfromafar 1mo agoAnd no monitoring of the proxy for unusual activity, either.
- irthomasthomas 1mo agoThey gave it a full package manager with internet access. They could have used a local cache and air gapped it, but they chose not too.
- Smaug123 1mo agoThey didn’t intend to give it Internet access. Artifactory is a caching proxy which can be scoped to specific package ecosystems, not a general Internet gateway (unless configured that way).
- shafir-info 1mo ago[dead]
- ifwinterco 1mo agoThey were using what (allegedly) they considered an incredibly dangerous “cyber grade” model - finding bugs in software is exactly what you would expect it to do
- tantalor 1mo ago> How were they supposed to know about "previously unknown vulnerabilities" Very simply, there is no such thing as bug-free software.
- ben_w 1mo agoIf this is your standard, I challenge you to name one currently operating business that isn't criminally negligent. I'm sure there's tens to hundreds of millions of them amongst the 37% of the world with no internet connection, but actually finding them listed on the internet will be somewhat of a challenge.
- hilariously 1mo agoNo, the entire marketing campaign for these models is "it automatically has godlike powers to exploit almost any software" - if you don't air gap such a capability you are inherently allowing shit to go down, any other interpretation is "OAI folks are too stupid to design a proper test".
- ben_w 1mo ago> "it automatically has godlike powers to exploit almost any software" No it isn't. Random people on sites like this mock them as if they're talking about having godlike powers. Each new model is "merely" a step up from what came before, the steps are frequent and rapidly improving, and just recently (in more than one AI company) crossed a threshold where that improvement made the tests dangerous. But even well before "godlike"*, there's plenty of research about how to cross air gaps. > "OAI folks are too stupid to design a proper test". Such binary thinking. It's very easy to say things are "obvious" after the fact. People do that all the time, e.g. how the Bay Of Pigs invasion was never going to work, or like the Zune wasn't a good product-market fit. Oh the stories I could tell if not for the NDAs. * whatever that's supposed to mean: https://news.ycombinator.com/item?id=40874779 https://news.ycombinator.com/item?id=40874779
- mcmcmc 1mo agoWho said criminally negligent? If the whole point is testing its exploitation capabilities and you don’t want it exploiting the environment to gain internet access, that’s why you air gap, to remove the possibility
- ForHackernews 1mo ago>How were they supposed to know about "previously unknown vulnerabilities"? You don't. That's why you unplug the Ethernet cable.
- ben_w 1mo agoHave you done that to your own machines? Seriously. If your reaction to the inability to know about previously unknown vulnerabilities is "unplug the Ethernet cable", why are you not doing that (and equivalent) right now to your phone, laptop, etc.? Remember, the open weights models are only a few months behind the private ones, so these events being from a few months ago means the threat of such models is something you ought to take with the same degree of seriousness that various commenters here deride OpenAI for not having had.
- echoangle 1mo agoAm I running a new model with unknown capabilities without safeguards on my own machine and then prompt it to do determine cyber capabilities? You don’t need to be a genius to see how airgapping would be a simple and much safer measure than using a VM.
- ben_w 1mo agoYou're on the internet, your threat is everyone else running a new model with unknown capabilities without safeguards. In particular, all my last paragraph. I do offline backups, which get physically unplugged between sessions. Even that might not be enough.
- ForHackernews 1mo agoThis is such a goofy comment. In your mind, there's no difference between the precautions a BSL-4 virology lab should take when working with an unknown pathogen and the precautions that literally everyone else in the world should be expected to adhere to? Because, hey, after they deliberately unleash their new unknown virus on the world, we're all going to face that same threat, right?
- kmeisthax 1mo ago> How were they supposed to know about "previously unknown vulnerabilities"? By disabling the models' own internal restrictions (or training without them) OpenAI was, effectively, running an AI malware lab. The standard IT practice for a malware lab is to airgap and wipe EVERYTHING, and to assume any software sandboxing is made of cardboard and niceties. You don't have to know about specific vulnerabilities to infer that they might exist, and there's defense strategies for unknown vulnerabilities. If a model found a way to jump an airgap by, say, using their CPU's clock generator like a Wi-Fi antenna, then yeah, that would be a "previously unknown vulnerability" and one that couldn't be reasonably foreseen. But it's reasonably foreseeable that a model with unknown cyber capabilities might figure out how to break out of a sandbox, given that sandboxes get broken out of all the time in security research. What I would have expected from a competent AI malware lab would have been, say, an inference box with a bunch of serial cables to individual blade servers with no network access and a preloaded drive full of Linux ISOs the model can stand up. When a model's context is wiped so is their attendant box, preferably by someone yanking the drive out and imaging it from a dedicated imaging machine. I can foresee other attacks (e.g. firmware persistence) that could have more exotic countermeasures designed for them, but this would at least be the bare minimum for taking AI safety seriously. (Y'know, the whole reason why OpenAI stopped being Open?)
- elpatokamo 1mo agoI agree with your first sentence, but not the second. Let's remember Hanlon's Razor. This would be a wild thing to do as a marketing stunt. They're essentially admitting to violations of the CFAA and are lucky Huggingface was sorta chill about the incident. My assessment? They deprioritized good cybersecurity controls in the name of moving fast. They had a single Artifactory instance shared across many (or all?) their training environments. And then, after the agents found a way to exploit it, they rebuilt Artifactory again and still set it up with one shared instance. That was careless, perhaps even reckless.