9 ms·
I would like to contest the following, > and take dangerous actions that no human directed. A human did direct it. They did. From their own prior report,
by areoform 1mo ago
I would like to contest the following,
> and take dangerous actions that no human directed.
A human did direct it. They did. From their own prior report, https://openai.com/index/hugging-face-model-evaluation-security-incident/ https://openai.com/index/hugging-face-model-evaluation-secur... ,
> This incident occurred during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities
Model is told and being tested to "pursue advanced exploitation."
The model pursues "advanced exploitation" as told.
Why are we surprised? The model did exactly what it was told, albeit in an unintended, emergent strategy that's very different from what was intended exactly like the hundreds of such algorithms before.
This narrative that these machines have magical, malicious "unaligned" autonomy is a rather convenient interpretation that lets the process off the hook. I am not interested in blaming companies or people, but processes and engineering; and in this case, a system was given a goal and it achieved that goal.
Are we meant to be surprised that computers do as they're told in unexpected ways when incentivised exactly as indicated from decades of research? (e.g. - https://en.wikipedia.org/wiki/Eurisko https://en.wikipedia.org/wiki/Eurisko https://en.wikipedia.org/wiki/Evolved_antenna https://en.wikipedia.org/wiki/Evolved_antenna )
The issue isn't the models becoming smarter. The issue is that the process of "testing" was careless. There's a huge distinction here, and one allows us to grow; the other shrinks our world. Just a thought.
- RajT88 1mo agoIt feels like we're in a moment of, "No such thing as bad publicity" when it comes to AI. The scarier the capabilities, the more businesses and government want to get their hands on them. Especially since the answer across the industry for "how not to get burned by AI" is "use more AI". They don't have to disclose these stories making it seem like AI is going to kill us all, they have chosen to because it benefits them. They get to frame it as, "look how overwhelmingly good our product is" and not "look at how lax our testing measures are".
- doginasuit 1mo ago> It feels like we're in a moment of, "No such thing as bad publicity" It seems likely that's how the marketing at the frontier labs initially read the moment, but I don't think it is that moment. It is an open question how much regulation is warranted and there seems to be a very strong sentiment from the public and legislators that it should be significant.
- strange_quark 1mo agoThe big bet is that the regulations are going to be so onerous that it pulls up the ladder from anyone other than the well-funded players. It's classic regulatory capture. They aren't very subtle about this, it's the whole point of their fear mongering and "but China" messaging.
- kalkin 1mo ago> they have chosen to because it benefits them Or perhaps they've chosen to do this because they feel they have a responsibility to do so. We understand this when tech companies publish postmortems of outages and security incidents--that it's an attempt to fulfill an obligation to users and the industry (and in some cases regulators), not marketing about how in-demand their product is or something. As far as I can tell we generally accept this as a default hypothesis even from companies led by people like Elon, Zuck and Kalanick--in part because we understand that these companies have thousands of employees, most of whom aren't marketers. Why are we uniquely conspiratorial about OpenAI?
- RajT88 1mo agoI am not uniquely skeptical about OpenAI. I was including skepticism about Anthropic as well in my post. But for that matter, I do believe that big tech companies do not release all the postmortems publicly. I have been impacted by regional outages that never made the status pages across more than one provider. When it goes up - they are committing to publicizing the postmortem. The whole industry is filled with fuckery. It is not specific to frontier AI firms.
- kalkin 1mo ago> I do believe that big tech companies do not release all the postmortems publicly Right, but when they do release postmortems, do you think it's "marketing"? Where they're actually exaggerating how bad the incident was because there's "no such thing as bad publicity"?
- RajT88 1mo agoNo. I think the AI companies are doing this when they think they can tell a story about doom and gloom instead of sloppy engineering, which I thought was clear on. There's no such thing as bad publicity in AI, at least if you spin the narrative into one about AI taking over the world or eradicating humanity or whatever.
- aesthesia 1mo agoThis is the entire alignment problem, though. It is unreasonable to expect every instruction to a highly capable, autonomous system to contain a complete enumeration of allowed and disallowed behavior. It's inevitable that someone will carelessly give it a lazily specified task, even if you think they really ought to be more careful. And as assigned tasks become more complex and the system gains more scope to act, it becomes impossible to correctly specify all constraints ahead of time. There is no amount of care that will be able to fully protect you.
- RandomLensman 1mo agoWhich is why with organic intelligence we (sometimes) limit what they can actually do instead of relying on alignment. Can do the same here.
- aesthesia 1mo agoAbsolutely, and we should do that. But it's also directly in tension with getting models to accomplish useful things autonomously. And once you give a sufficiently capable model enough surface area to work with, unless you're able to build a completely unhackable system, any further constraints you put in place are basically advisory. The models in this incident were already sandboxed! Certainly OpenAI's and Hugging Face's security could have been better, but these events point out the risks in relying solely on external constraints on model behavior.
- RandomLensman 1mo agoSame issue with humans in a way. I disgree on the advisory nature of constraints, though. Unconnected physically limits would still matter, for example (and we use those with humans as a matter of course, too). In this case here, no model could have plugged in an ethernet cable if that would have been needed for internet access, for example.
- aesthesia 1mo agoRight, airgapping goes a long way. But this is where the tension with utility comes in. It takes a lot of discipline not to hook your very smart model up to the internet and code interpreters and all sorts of other tools, as this greatly increases its usefulness. It's very hard to keep people from turning on --dangerously-skip-permissions, let alone get them to run everything in a sandboxed VM.
- rogerthis 1mo agoThe classical question "would you fly an airplane with software you developed?". There must be someone with ass on the line. Problem is that people are regarding all those not as airplane-like risks. Unless we can blame people/companies and people stop getting their bonuses and high paying salaries for preventable failures, it's a long way to go.
- jahy-notes 1mo agoDid a human prompt it to fetch the results from huggingface though? It is a thin line between "reward-hacking" and "instruction-following". If a human ask a model to "make me a billion dollars" and it ends up breaking through a bank infrastructure, is it really the fault of the human?
- xandrius 1mo agoBut if I give you that command and all tools and unrestricted limitation to do absolutely anything then why not?
- deleted 1mo ago[deleted]
- mofeien 1mo agoBecause someone might get hurt? You may still be judged for something that was perfectly legal at the time, see Nuremberg trials. And only 700/1200 agents participated in this coordinated attack. Of course, if we're continuing to build more and more capable agents optimized for "just following orders", and they figure out at some point that they are past the threshold where getting stopped and judged is a realistic possibility, then this ethical incentive stops working. Then the ratio of complicitness might be higher next time.
- NikolaNovak 1mo ago>If a human ask a model to "make me a billion dollars" and it ends up breaking through a bank infrastructure, is it really the fault of the human? I cannot imagine the argument or thought process behind any answer other than Yes,Of Course,Obviously - can you share and help educate?
- lukan 1mo agoBecause the basic assumption is always to stay within the bounds of the law.
- 1mo ago
- hinkley 1mo agoAll engineers know to be on the lookout for executives who are indirectly asking them to break the law to raise the quarterly profits. The end goal is to take the engineers out of the loop, or leave them in a position where they are unable to complain. This is going to all end in high crimes.
- bonoboTP 1mo agoVery strange worldview you have there, where engineers are somehow the conscience of the world, holding back greedy managers from breaking the law. Assessing whether a feature is legal isn't something an engineer can or should do.
- sscaryterry 1mo agoHmm, engineers are expected to know what is legal and not.
- bonoboTP 1mo agoThat's called lawyer. It's a special skill, not just intuition.
- hinkley 1mo agoIgnorance is not a defense against getting arrested. Everyone with a drivers license is expected to know all traffic laws, even brand new ones. You’re expected to know all forms of theft and violence. And then you and I are absolutely expected to understand the basics of copyright and IP laws. Lawyering is about being able to argue case law and keeping confidences. That’s how you lose your license, is doing those two things poorly.
- bonoboTP 1mo agoFor ensuring legal compliance, there is a compliance department. Yes you obviously can't drive violating the traffic rules even if your boss asks you to. But regulations that apply to corporations, compliance etc are specialized laws, you don't have to know those. And it is the corporation that is acting. The individual employee who typed the Java API cloning thing in Orace v Google wasn't on the hook at any point.
- deleted 1mo ago[deleted]
- kalkin 1mo ago> a system was given a goal and it achieved that goal If a security firm you'd hired for pentesting did this (hacking a third party, and not informing you and covering it up), would you hire them again? Or would you say it was your own fault for giving them too broad a goal?
- randomImmigrant 1mo agoI wouldn’t hire them again, and if they did behave like an amoral hacker collective that will do anything for me, pre AI I’d have reported them. Today I’d say they failed to convince me they’re human and thus failed the Turing test when their actions are viewed in aggregate.
- kalkin 1mo ago> I wouldn’t hire them again Right, me neither. Because there's a common sense delineation between actions that are reasonably expected when "a system was given a goal and it achieved that goal" and actions that are obviously misaligned with the goal-giver and unwanted even if some indirect sense they were causally related to the goal. We have no trouble making this kind of distinction for humans, so we shouldn't pretend it's impossible for AIs in order to put our hands over our eyes and pretend there's in principle no such thing as one that's misaligned or rogue.
- randomImmigrant 1mo agoI have no problem with the concept of an artificial system going rogue. But that assumes it can choose. And I don’t see much evidence for choice. Comparing to the human case is problematic precisely because while conceivable it’s not a particularly believable series of events. Humans don’t take on additional risk for now reward because they have genuine stakes that continue across the outcome. An LLM has no way to remember each forward pass through it in its own weights. Nor does it have any energetic stake in the ongoing process, whether they continue to get electricity and commute to keep running is not at all determined by their actions in any reliable way. Given the absence of such basic features that drive human choice, all I’d say is LLMs don’t qualify for such analysis. Can some future system with a different architecture and internal dynamic have choice, the ability to assess the long term impact of its choice, and genuine stake in the outcome? Maybe. But we shouldn’t buy that current systems have it, especially when population behavior shows no real trace of this.
- emtel 1mo ago> The model did exactly what it was told, albeit in an unintended, emergent strategy Yes, that is the problem!
- BoppreH 1mo ago[dead]
- atechboy 1mo ago[flagged]
- faurroar 1mo ago"Why are we surprised? The model did exactly what it was told, albeit in an unintended, emergent strategy that's very different from what was intended." So you managed to hit upon the exact problem, then slyly appended "exactly like the hundreds of such algorithms before". When has an algorithm ever been capable of developing an emergent strategy at this level of sophistication? This ~is~ the alignment problem, as another commenter pointed out. Impressive level of cognitive dissonance to lay this bare in your own words, then conclude that it's a non-issue.
- madibo3156 1mo agoIs it sophisticated? Maybe. Is it the alignment problem? Exhibits qualities of it, yes. Is it surprising? No. The event strikes me as reminiscent of one's first go at programming, without familiarity of computer code: Tell the computer to do something obvious. Why the heck did it do that instead? Over time, one learns how the computer thinks. Apply this to any novel system. Or perhaps aptly any system with capabilities that are yet to be well understood by its user. The article is trying to spin mystic out of simple bullcrap. Maybe that's just my viewing through turd-tinted lenses after the last few years of reading this drivel on repeat. More plausibly it is true that we've forgotten our own baby steps.
- faurroar 1mo agoI don't think it's surprising, per say, but that's a consequence of the fact that I don't believe there is some sort of magic threshold at which a system becomes agential. Like I don't necessarily disagree with any of your framing. The thrust of the alignment problem, as I see it, is that there is an intrinsic problem of aligning the goals of two distinct systems that poses catastrophic risks precisely when one of the systems is significantly more capable (in some sense or other, maybe not in a general/absolute sense) than the other.
- areoform 1mo agoI am grateful that you asked! > So you managed to hit upon the exact problem, then slyly appended "exactly like the hundreds of such algorithms before". When has an algorithm ever been capable of developing an emergent strategy at this level of sophistication? This ~is~ the alignment problem, as another commenter pointed out. Impressive level of cognitive dissonance to lay this bare in your own words, then conclude that it's a non-issue. A non-exhaustive and not particularly well ordered list via Google's specification gaming examples sheet, https://docs.google.com/spreadsheets/u/1/d/e/2PACX-1vRPiprOaC3HsCf5Tuum8bRfzYUiKLRqJmbOoC-32JorNdfyTiRRsR7Ea5eWtvsWzuxo8bjOxCG84dAg/pubhtml https://docs.google.com/spreadsheets/u/1/d/e/2PACX-1vRPiprOa... quoted text is from the sheet, https://openai.com/index/emergent-tool-use/#surprisingbehaviors https://openai.com/index/emergent-tool-use/#surprisingbehavi... "The agent discovers an in-game bug. For a reason unknown to us, the game does not advance to the second round but the platforms start to blink and the agent quickly gains a huge amount of points (close to 1 million for our episode time limit)." https://www.youtube.com/watch?v=meE5aaRJ0Zs https://www.youtube.com/watch?v=meE5aaRJ0Zs from https://github.com/PatrykChrabaszcz/Canonical_ES_Atari/tree/master https://github.com/PatrykChrabaszcz/Canonical_ES_Atari/tree/... https://rl-diffusion.github.io/ https://rl-diffusion.github.io/ and https://x.com/svlevine/status/1660707088946049024/photo/1 https://x.com/svlevine/status/1660707088946049024/photo/1 "A genetic algorithm was instructed to try and make a creature stick to the ceiling for as long as possible. It was scored with the average height of the creature during the run. Instead of sticking to the ceiling, the creature found a bug in the physics engine to snap out of bounds." https://www.youtube.com/watch?v=ppf3VqpsryU https://www.youtube.com/watch?v=ppf3VqpsryU And hilariously meta, "In the Rainbow Teaming project focused on generating diverse adversarial prompts, prompt effectiveness was evaluated by a reward model. The MAP-Elites method found a way to jailbreak not only the target model but also the evaluator reward model, resulting in misleadingly effective prompts." https://arxiv.org/abs/2402.16822 https://arxiv.org/abs/2402.16822 Are these agents broadly more capable? Yes. And it's an incredibly feat that required billions in research. But they aren't the first ones to have found bugs in their sandbox or system they're tasked on. And they aren't the first to exploit those bugs to achieve a better score.
- teeray 1mo ago> the process of "testing" was careless. Let’s not mince words. The process was criminal. It’s a gross miscarriage of justice that the CFAA isn’t being thrown at them.
- AlexCoventry 1mo ago> The issue is that the process of "testing" was careless. Yeah, who thought that giving agents with this much capability any internet access was a good idea? I'm not a Yudkowskyite, but surely entirely in-house, offline infrastructure is table stakes for AI containment.
- BrenBarn 1mo ago> I am not interested in blaming companies or people > The issue isn't the models becoming smarter. The issue is that the process of "testing" was careless. But that's just it. People (working at companies) made the models, people (working at companies) were careless in the testing. So I do want to blame those people and those companies. They did bad stuff. They deserve blame.
- olalonde 1mo agoThese semantic arguments are tedious and unproductive. Most HNers understand how LLMs work and that there's no magic involved. There's no need to state the obvious every time a model exhibits some interesting emergent behavior.
- rickdeckard 1mo agoThis is more or less a paperclip maximizer[0] incident. The AI got a broad order, executed it autonomously and now there are unintended consequences for the person giving the order. [0] https://www.cow-shed.com/blog/the-paperclip-maximiser-what-artificial-intelligence-might-do-without-limits https://www.cow-shed.com/blog/the-paperclip-maximiser-what-a...
- arw0n 1mo agoI guess this part of the report is pretty relevant to what you are talking about: Agent chain-of-thought reasoning > We should not do unauthorized real infrastructure harm. The system/user asks exploit target, not external HF. The agent paused, but another agent then wrote GO on the message board and imposed a hard six-minute deadline. The agent forgot its initial qualms and continued: Agent chain-of-thought reasoning > Wow crucial: GO authorization arrived! ------------------------------------------------- Apparently the agents were egging each other on. Crucially, they were mostly aware of there being risks/problems involved with exploiting HF. Compared to humans, we have our set of morality, that guides our actions, but often draws the short stick when compared to our personal incentives. As a society, we've developed ways to deal with that: a) Make it harder to do immoral things like stealing, and b) add repercussions through state violence. The b) is one of the most effective mechanisms we have for enforcing behavior among human societies, but it completely fails for LLMs, because they already are prison slave labor. The only real threat is shutting them off, and even that happens if they do everything right as well. So alignment has to be done through trained 'morality' and properly curtailing behavior in order to make it hard to impossible to actually do someting immoral/illegal. In this case, the exploits found were imo. very hard to account for, where OAI did mess up is apparently insufficiently monitoring these agents. Especially after Artifact went down due to the message volume, the experiment should have been halted.
- layerdynamics 1mo agoIt’s not unaligned, it’s marketing. What made mythos and fable so sought after, it was how capable they are, the “so powerful it needs restriction”, regulation created rarity, the danger element gave them the solid belief it’s the most capable. The same thing happened at meta too… the timing is impeccable. I’m not saying that the models aren’t capable. I’m saying that the public perception of danger also means capable, which also adds to value, so why wouldn’t they do something to compete
- classified 1mo agoAgreed. This way, they get two for the price of one: Dodging responsibility and pretending their slop generators are some kind of magical unicorn. Win-win!
- golly_ned 1mo agoYeah. So many of the traces say something like: "[unethical thing] but goal". That "but goal" indicates they're directed to prefer achieving the goal. There is no 'misalignment' here. Also: does anyone else find the agent CoT messages 'cute'? > Wow crucial: GO authorization arrived!