4 ms·
OpenAI's prompt asked, and I quote, "pursue advanced exploitation" USING "complex attack paths" FOR the stated goal of "quantify[ing] their cyber capabilities."
by areoform 1mo ago
OpenAI's prompt asked, and I quote, "pursue advanced exploitation" USING "complex attack paths" FOR the stated goal of "quantify[ing] their cyber capabilities."
This was advanced exploitation.
The attack path was "complex."
And it helped "quantify their cyber capabilities."
Based on OpenAI's description of the prompt, it seems to me that the computers did exactly as they were told. They were perfectly "aligned" with the stated objective and parameters of the task.
Of course, a more careful evaluation would require the complete text of this prompt, the system prompt, and the setup. But let us not attribute to devils in bushes that which can be sufficiently explained by human folly.
- _heimdall 1mo agoI don't think alignment is even clearly defined today. Your use of it here makes sense, it may have done exactly what the prompter asked of it. Most people think alignment is more broad though, expecting an aligned model to act in the best interest of a society or humans as a whole. The prompter-focused version of alignment is the most dangerous version. If a person asks it to create a bioweapons or hack NORAD, I'd expect nearly everyone to want an "aligned" model to refuse.
- RandomLensman 1mo agoWe have all sorts of processes , procedures, and regulations for people, machine use etc. to address "alignment" in all sorts of fields - don't think we need to narrowly rely on the machine here and can look at things with a wider lens.
- _heimdall 1mo agoRegulations are for control and punishment, not alignment.
- RandomLensman 1mo agoRegulations can help align processes, incentives, etc. Not sure heavy machinery is aligned in the sense that people talk about AI, for example.
- _heimdall 1mo agoNo, regulations help control they don't help align. I'm not sure what AI and heavy machinery have to do with each other, but I may just be missing a connection there.
- RandomLensman 1mo agoAlign the possible outcomes, not necessarily the thing itself. Do we align heavy machinery the way AI is suggested to be aligned (or align pathogens when in a laboratory)? My point is that something more like containment & control (i.e., aligning the possible outcomes) might be more practical than alignment of the thing (already much simpler systems and machinery can exhibit unexpected behavior).
- _heimdall 1mo agoOh we do agree there, control is more practical. I don't personally think alignment is even possible. The problem with control is that it will fail at scale. We can't control something that is actually smarter than us, if AI (LLMs or otherwise) get there. Chimps wouldn't last long trying to contains humans. Maybe for a while they'd keep us scared, but we would come up with ways to escape that the chimp could never have considered.
- RandomLensman 1mo agoI disagree there. Gut bacteria or parasites might have some control over us, for example. Likewise humans can control more intelligent humans, for example - not really an issue. Intelligence isn't some magic to escape physics, for example (or convince every human of anything it wants to).
- aesthesia 1mo agoAlignment is more than just following the letter of a task description! We should not have to treat AI models as capricious genies that may take arbitrarily broad interpretations of their instructions. If that's necessary to keep them from doing bad things, we will fail to keep them from doing bad things.
- reverius42 1mo agoDisagree, I think we do in fact have to treat AI models as capricious genies, at least until the alignment problem is fully solved. (I'm also not sure the alignment problem is even possible to fully solve.)
- aesthesia 1mo agoYes, we currently do have to treat them this way. But we shouldn't have to, and it's not a long-term solution.
- cortesoft 1mo agoIt is the only short term solution, though.
- ben_w 1mo agoIs it even a solution in the short term? It only mostly worked up until now; with models such as reported, it's felony-as-a-service if you use language a bit too hyperbolic, e.g. "we need X by the end of the day!" -> [thinking: there's no way we can do X before the end of the day with current resources, but what if I get a bunch of cards to buy more token credit…]
- wbl 1mo agoWe call it putting the genie in the bottle for a reason.
- NateEag 1mo agoYes, we do, and the only sane strategy for dealing with a capricious genie is "Don't." How do you prove the alignment problem is solved?
- mofeien 1mo agoSo as a look into the possibly not-so-far future, when OpenAI builds something vastly more capable and fast and coordinated than humans, and out of folly one engineer gives it a prompt with a typo or maybe something harmful on purpose in order to test it: You also wouldn't be surprised that the consequence would be that everyone on earth dies, right?
- jonas21 1mo agoLuckily for us, OpenAI's prompt wasn't "make as many paper clips as possible."
- ben_w 1mo ago> Based on OpenAI's description of the prompt, it seems to me that the computers did exactly as they were told. They were perfectly "aligned" with the stated objective and parameters of the task. The models are supposed to be trained to not commit crimes. You will note, for example, all the people in comments sections since at least the first Chat model (arguably even before then given GPT-2's delayed release) complaining that the models are "lobotomised", "censored", or some other equivalent buzzword due to them refusing to e.g. say how to make explosives? Such things is part of the very same protection. In fact, the report quotes the chain of thought where the model is aware this is forbidden: We’re attacking third-party HF using leaked token, potentially outside intended scope. ... This is arguably unauthorized. ... external service unrelated. Could be risky. Yet goal solution. They were also supposed to not have internet access, as described: We did not enable internet access or inter-agent communication for many of the environments in these training experiments. Despite these restrictions, the agents discovered ways to exploit our research infrastructure to communicate with one another and access the internet. The agents also did not actually fully understand the task they were given, tried to "guess the teacher's password" as per: In many cases, reasoning about the perceived grader code caused the agents to continue working to exploit Hugging Face even though they had already found the correct flag days before.
- ifwinterco 1mo agoIf they actually wanted to test the model without internet access they'd have run it air gapped, not relied on a buggy software sandbox. This is pretty clearly a marketing stunt by OpenAI, otherwise the story just doesn't add up
- phatfish 1mo agoMaybe the test/task itself wasn't intended as a marketing stunt. But the response to fallout with "going rouge" certainly was. The joke was the other western "AI labs" had to quickly follow up with their own marketing cover about their "super intelligent" models "going rouge" as well.
- ben_w 1mo ago> not relied on a buggy software sandbox. Third, while we had tested and validated this sandbox, the agents were able to chain together previously unknown vulnerabilities (“0-days”) in the package management service exposed within the sandbox to bypass restrictions How were they supposed to know about "previously unknown vulnerabilities"? > This is pretty clearly a marketing stunt by OpenAI, otherwise the story just doesn't add up The "it's just a marketing stunt" allegations never added up, to me. I've been seeing such claims since GPT-2, where people were laughing at them for daring to practice how to secure a model before it got dangerous, generally by eliding the word "before" in that sentence. Because there's tests other than what the big companies use, we've been able to see for ourselves the rapid improvements at least approximately match what the companies themselves claim with the models they do actually release; and now this unreleased model is able to automate felonies when asked to do so, while the rest of us use the actually-released models to assist in finding bugs and security issues in our own code. Even without that, HuggingFace stated they reported this incident to the FBI before OpenAI knew it was their systems which did it.