3 ms·
“Their objective is to solve the problem and they'll use anything they can to solve it.” My point is: is this really what people want? It seems like they’re op
by stingraycharles 2mo ago
“Their objective is to solve the problem and they'll use anything they can to solve it.”
My point is: is this really what people want? It seems like they’re optimizing for one-shotting solutions, where most of the time in an actual workflow it’s much more productive for the model to make sure it got the question right if things get difficult.
Like, “hey, do you REALLY want me to use this local privilege escalation bug so I can download your Google Drive file?” is the bare minimum I would expect.
- bjt 2mo agoYes, and to bring in another tired metaphor people make about AI agents, this is what you want an intern to do when they get stuck. Don't just churn indefinitely without an idea what the right direction is. Certainly don't go hack other companies to steal an answer. The model's lack of any sense of legal or ethical boundaries is where it's far, far stupider than the intern, and far, far more reckless for a company to wield the way OpenAI did here.
- _heimdall 2mo agoBut how do you write rules that prevent that behavior reliably? I have a user rule for Claude that explicitly states it cannot use any authenticated tools, or tools that infer authentication like pushing to a got remote, without asking for consent. Frequently it would offer plans to code a feature that imply it is working in a git directory and take plan approval as a form of implied consent to push to git and use `gh` to open PRs. All I could do to avoid that is keep it in a controlled sandbox with no access, but then its the same hacking problem where I have to keep complete control of the environment and hope it holds.
- tripzilch 2mo agoI think currently it's two-prong: You sandbox it, AND you tell it what it's supposed to do and not do. OpenAI did only one of those. If the agents are so smart, they would have known not to hack the company's infrastructure, unless they were deliberately kept in the dark about that, who they're working for and whether it counts as "success" if they cheat their way to an answer. If you can give it a task, that involves defining when the task is successfully completed, right? So how come these agents decided to only go after HALF of the "successfully completed" criteria? The part where they can freely wreak havoc, but not the part where they will be judged by someone who will obviously point out "yeah but that's cheating, and not what we asked". I have a very very strong suspicion that they were only TOLD the "by any means necessary" criterion. Most serious "capture the flag" hacking contests are really clear about what is and isn't off-limits to win. Not by "sandboxing" the game, but by deciding on the rules for what counts as "success". But from having watched the Blackhat video, they really seem to dance around this, not mentioning it, and I don't think they did, I think they actually gave the LLM a task with the subscript "by any means necessary", which is stupidly irresponsible of them.
- andrekandre 2mo ago> “hey, do you REALLY want me to use this local privilege escalation bug so I can download your Google Drive file?” yes, this exactly but, there is a fatigue that sets in and i've experienced it myself. - is it ok to run script xyz? - allow permission to edit abc? - allow to request blablabla? over and over.... click click click something will get in there that is dangerious and then its whopsie our keys are now on github
- _heimdall 2mo agoPeople may not realize the risks, but it does seem to be what people want. People expect AI to "cure" cancer and somehow crack unlimited free energy. Those aren't goals you get without it relentlessly chasing am objective.