3 ms·
There are two things I don't understand about this story. First, why does an agent get any write access to artifactory at all? Second, why is the artifactory
by choeger 1mo ago
There are two things I don't understand about this story.
First, why does an agent get any write access to artifactory at all?
Second, why is the artifactory cache not disconnected from the net? Surely you'd not feed it with new software versions while the eval or training is running.
- 1dom 1mo agoFrom what I can understand from reading a few different, slightly conflicting, versions of these events: they weren't given write access. They found a zero day exploit that allowed them to create folders, and the folder names were initially used for agents to communicate. I'm not sure artifactory was connected to the net. Some agent sandboxes had internet access and were able to communicate with ones without access via artifactory.
- choeger 1mo agoI read the agents used SSRF via artifactory to gain uncontrolled access to the net. Apparently their intended net access went through a tightly controlled proxy. Even that appears to be very risky, tbh. If I was to setup a sandbox for such a complex and autonomous system, I'd probably point them to an archive-like cache for net access and cut their comms at the package level.
- izend 1mo agoWhy wasn't the traffic in/out of the boxes that the agents were running on monitored?
- consumer451 1mo ago> Why wasn't the traffic in/out of the boxes that the agents were running on monitored? I have to assume: move fast and break things. I don't mean this to be taken as a hot take. The startup scene loves to poo-poo on things like this as unnecessary overhead. OpenAI and many others like to operate as a startup, to move fast. Disclaimer: in far, far lower-stakes situations, I certainly do this myself.
- pixl97 1mo agoI have a few 'conspiracy' theories on this that go from likely to sci-fi. My two big ones for this would be 1. They do monitor the AIs attempting to hack but for different reasons than you expect. Instead of making models that don't hack they are trying to build the most efficient hackers in the world and sell this capabilities to governments for billions. Because of this they generate terabytes of hack attempt logs and agent history doing this hacking. So when a new model came out with better abilities what they were looking at changed and they didn't realize it. They were already numb to alarms and missed when the danger occurred. 2. Like the above, they generate terabytes of logs per day. Because there is so much data AI filters and monitors almost all of it flagging things that a human should review. But for some reason this model didn't set off those flags. The protection model classified this behavior as perfectly safe. Number 2 sounds kind of like a sci-fi conspiracy but it seems that almost all models judge content generated by the same model or family of models as 'better'. It's predicted that models in a judging context could allow things to slip by as an emergent behavior of reading the text.
- sensanaty 1mo agoBecause they're incompetent or simply don't give a shit.
- lovich 1mo agoif you wanted to sandbox their access to the internet, why give them any physical access at all?