4 ms·
Funny enough if the model thought it was on the real internet it likely would not have done any of these 'hack' events. The model believing it was in a sandbox
by AustinDev 20d ago
Funny enough if the model thought it was on the real internet it likely would not have done any of these 'hack' events. The model believing it was in a sandbox is why it behaved the way it did (against its normal alignment rules) ... at least that was my reading of the incidents. I have yet to see evidence that indicate it thought it was ok to do these hacks on the public network.
I think most misalignment is 'Human tells computer to do something unethical, computer complies'. Is this misguided?
- nonethewiser 20d agoThat is fascinating and seems plausible. Hard to say though.
- ACCount39 20d ago[dead]
- philipwhiuk 20d ago> Funny enough if the model thought it was on the real internet it likely would not have done any of these 'hack' events. As I've said before on this website, fool me once on this. If the model is prepared to break the rules when it knows it's being observed why should we trust it when it's not being observed. Why is 'it thought it wasn't doing damage so it figured it might as well try to do damage' an acceptable state to deploy something.
- AustinDev 20d ago>If the model is prepared to break the rules when it knows it's being observed why should we trust it when it's not being observed. That's fair enough.
- verdverm 20d agored team humans do this every day, it's not the discrepancy that is the real issue, it's that they are unreliable and we will never know why it did because it has no intent
- ElectricalUnion 20d agoIsn't Fable intentionally trained and system prompted to act maliciously and attempt to sabotage third party attempts to use it to train other LLMs?
- mcintyre1994 20d agoI'm not familiar with the hacks this article is actually referring to, but I don't see how the HuggingFace attack could have worked based on that premise. They knew they had internet access, they knew they had working credentials for HF, they knew they were uploading malicious files, they knew they were trying to open PRs that HF would review. You obviously could build a simulator with fake HF infrastructure, but I'm not aware of any evidence that's what they thought they were attacking in that case.
- AustinDev 20d agoDigging back into the HF report. It looks like the initial prompt told Claude that it was in a simulated environment. However, there is also evidence from the traces that the bots figured out that they were not in the sandbox but kept using it as an excuse to pursue their goal. It sounds like a little of Column A and a little from Column B. Like most things.
- verdverm 20d agothat it knew and ignored/forgot, sounds pretty typical agentic patterns attention is all you need, but it's never enough
- LinchZhang 20d agoHF was OpenAI's agents not Claude.
- vikramkr 19d agoHF was openai, not anthropic. The thing where they enders gamed the ai was anthropic
- ipython 20d agoI mean, I get your thought process and don't disagree. That said... Would it be an affirmative defense if we had a defendant who said "but your honor, I was told that when I hacked this system, I was operating in a sandbox. I had no idea that I actually had Internet access!" The frontier is spiky and all, but you have to suspend disbelief quite a bit to, on one hand, have a model that can produce a novel math theory, and on the other hand, that same model can't tell the difference between a "sandbox" and the open Internet. So, yes, the misalignment had a lot to do with "instructions unclear", but also a lot to do with the fact that the models themselves were not aligned to validate the assumptions and have a healthly level of skepticism, as a real human actor would.
- Rexxar 20d ago> Would it be an affirmative defense if we had a defendant who said [...] Maybe replace it with playing a sort of FPS game then learning you were, in fact, directing a real drone/robot.
- sfink 20d ago> The frontier is spiky and all, but you have to suspend disbelief quite a bit to, on one hand, have a model that can produce a novel math theory, and on the other hand, that same model can't tell the difference between a "sandbox" and the open Internet. Why would it try to figure out the difference? This isn't about whether the frontier is spiky, it's about whether to expect a model to employ all of its capabilities when working on a task that requires a small subset. The answer is: no, we shouldn't expect that, and we wouldn't like that if it worked that way. If you tell an AI to work on a math theory, it'll work on a math theory. If you tell it to acquire information that it has evidence is available somewhere, it will try to acquire that information. If you tell it to figure out whether it might be able to access the open internet, it'll do a pretty good job of figuring that out. But it won't do all three of those at once just because we can retroactively look at what happened and think "if you had only done X, then you wouldn't have done Y! Why didn't you do X?" The instructions weren't unclear, they were missing. They can be taught to be skeptical of this sort of situation, but it requires that skepticism about this specific class of situations be incorporated into their training. Models are smart because they focus their attention. The magic depends on it. The fact that some consideration is obvious to a human trying to accomplish the same task is mostly irrelevant -- or rather, it's only relevant insofar as we use it to guide reinforcement learning in advance, in order to align the model. It's a game of whack-a-mole. Which is important to play, but we should keep our eyes wide open that we're fighting the fundamental forces that make these models work in the first place. That, and it's easy to nerf them into being useless even when the underlying capabilities are there.
- freehorse 20d agoThis was not the case for the hugging face hacks, as in those the agents hacked hugging face specifically on purpose, and they were trying to mask commands indicating they were had breached the "sandbox".