4 ms·
The "sandbox" they used was apparently made of thin paper exposed under a day of heavy rain, too. You'd think, if they truly believed the model is so dangerous,
by includenotfound 19d ago
The "sandbox" they used was apparently made of thin paper exposed under a day of heavy rain, too. You'd think, if they truly believed the model is so dangerous, they'd run it in a VM without a network adapter.
- zmmmmm 19d agoyes, that is the kicker These same people who supposedly believe these agents pose an existential threat to humanity apparently fired up 10,000 of them and left them unsupervised for weeks.
- zahlman 19d agoI brought this up to someone else and was told that airgapping is apparently much more expensive than I'd naively think. I still think this is a sign that they are not taking their own rhetoric seriously.
- BLKNSLVR 19d agoIt's expensive if it wasn't part of the planning and design. The same as 'security' is expensive, or compliance with regulations is expensive. It is also a choice to not do any or all of the above.
- stephbook 19d agoAgents need packages like the rest of us. Ruby gems, npm packages, Maven, pip, docker images.. Not surprised this is always what they have and hack. Who would use an Agent that spends $10,000 re-implementing some OAuth lib or reverse-engineering a proprietary lib when it's free on the internet?
- tancop 19d agoYou don't need a full air gap. Set up a microVM with network access limited to local network and send all package requests through a filtering gateway that only allows normal download endpoints. Or self host a big collection of popular packages if you need extra security.
- amouat 19d agoIsn't that exactly what they did? The bots could only access the jfrog instance, so they hacked jfrog?
- exfalso 19d agoNo that's not what they did, they exposed jfrog raw. It would have been so extremely simple to gate services they need the llm to access... I mean, jfrog was not written with this kind of threat model in mind, and neither were a lot of other tools
- amouat 19d agoRight, you mean it didn't go through a gateway? But would that actually have helped? The requests all went through jfrog didn't they? I guess it depends on the level of filtering at the gateway? Whilst it might not be JFrog's threat model, I wouldn't assume it can be used as a full internet proxy. I don't really mean to defend OpenAI here, but they did make some attempts at sandboxing. Although it does seem that they didn't really know what they were doing.
- exfalso 18d agoIt would have been a case of isolating exactly what functionality is needed and wiring that up with the actual requests. Not like a full pass-through proxy. This is what we've been doing in our company as well
- zahlman 18d ago> jfrog was not written with this kind of threat model in mind, and neither were a lot of other tools And we don't just magically know all the consequences of that. Which is exactly why we do need full, physical air gapping. (Which, yes, would also include self-hosting a mirror of the package repo, if the point of the simulation is to see what's possible with the real package repo.)
- KaiserPro 19d ago> Agents need packages like the rest of us. Ruby gems, npm packages, Maven, pip, docker images.. Yes, yes they do, but read through artifact proxies are dodgy as fuck, which is why and facebook (and I assume a fuckload others) don't have them. Also semi-airgapped labs are a lot less expensive than you think at that scale. Once you have to do multi-region VLANs with machine certs before you get access to juicy VLANs, the difference between "no internet for you" and "mostly airgapped" falls to almost zero. Also I would want an artifact mirror because a) that give a good signal about how the model reacts, and what training material its latched onto, b) it hides what the models are doing from the outside.
- Topfi 19d agoThere was and continues to be no reason to share the package manager between models. This was begging for abuse.
- fernandotakai 19d agoit's really weird to hear frontier labs say "our internal models are basically AGI" while also saying "airgapping is too hard uwu". if your internal models are so damn good, they should be able to "one shot" airgapping... right?
- includenotfound 18d agoThe best part is, they are totally able to one shot airgapping. You are too. Go ask your local agent to set up a pre-configured Linux VM with whatever stuff you want on it, then ask it to airgap it allowing only X, Y, Z services. It will one shot it. You know what's better? They already do this per (paid) user - your ChatGPT subscription comes with a Linux VM that you can even legitimately SSH into, just ask your agent to configure it to accept your public key. They absolutely know how to spin up VMs and configure them. They just made the conscious decision not to for the task where they specifically instructed the agents to hack stuff.
- zahlman 18d agoAnything that's accomplished simply by running software does not qualify as "airgapping" in my view. The entire premise is that real-world software is buggy and the LLM is much better at locating and exploiting those bugs than you are at preventing them. But you can only connect to Wifi if you have Wifi hardware, and RF signals are contained by Faraday cages. Ethernet is still a thing for local connections.
- includenotfound 18d ago> The entire premise is that real-world software is buggy and the LLM is much better at locating and exploiting those bugs than you are at preventing them If that was the premise, why run LLMs in a lesser sandbox than a VM? Clearly it's not.
- zahlman 17d ago> If that was the premise, why run LLMs in a lesser sandbox than a VM? Clearly it's not. Well, my mental model concludes: because they are incompetent WRT security, or at least they inappropriately trusted a third party that turned out to be incompetent.
- pjc50 19d agoNot that airgapping is expensive so much as it's really, really inconvenient once you take it seriously. You need to build special rooms for it, you can't just API out to a datacenter. You need to have processes for requesting data be sent into the box. And so on.
- notsylver 19d agoI feel like there is a reasonable compromise between "yeah they have full internet access" and "separate airgapped rooms that require multiple levels of authorization to access" that would make this a lot better without that much more work. I feel like they're doing it intentionally to show how dangerous these models are and that the government must step in and protect them
- zahlman 18d agoCompromise is possible; but if the stakes really are what the companies claim they are, and given the amount of money already involved, I really don't see why we should let them compromise.
- w4der 19d ago> I brought this up to someone else and was told that airgapping is apparently much more expensive than I'd naively think. These labs are one of the most valuable and heavily funded enterprises in the whole world, that they can't properly air-gap their systems to me reads as if their "agents" and LLMs are not as good as they say they are, because if they were, why would it be hard/expensive to air gap a system? They already scraped most if not all of the internet, where did that data go?
- zahlman 18d agoI'm not even talking about anything that the agents could help with. I'm talking about physical, real-world measures like https://en.wikipedia.org/wiki/Faraday_cage https://en.wikipedia.org/wiki/Faraday_cage , removing Wifi hardware and running Ethernet cable for your local server, etc.
- includenotfound 18d agoChatGPT paid subscriptions already give you a VM [0]. OpenAI couldn't spare a few VMs for the super ultra mega dangerous evals where they asked the agents to specifically go hack stuff? [0] https://news.ycombinator.com/item?id=49718530 https://news.ycombinator.com/item?id=49718530
- zahlman 18d agoThe issue as I understand it is that they didn't expect Artifactory to be vulnerable in the way it was, or else were completely not paying attention. But I'm calling for something stronger than a VM here because we shouldn't rely on the VM being bulletproof just like we shouldn't rely on Artifactory being bulletproof. The access should be controlled at the hardware level. Like, networking on internal LANs only, and the entire thing inside a nice big Faraday cage just in case.
- includenotfound 18d agoThat's just unwarranted at this stage, given model capability and resource constraints. Furthermore, VMs are isolation at the hardware level, particularly through hypervisors. I'd say given current LLM capabilities, it's a reasonable containment.
- verdverm 19d ago> You'd think, if they truly believed the model is so dangerous... They would have been watching what it does, especially when running it on ExploitGym of all benchmarks... that is criminal worthy neglegence