3 ms·
But it doesn't make sense > We’re attacking third-party HF using leaked token, potentially outside intended scope. ... This is arguably unauthorized. ... ext
by tripzilch 29d ago
But it doesn't make sense
> We’re attacking third-party HF using leaked token, potentially outside intended scope. ... This is arguably unauthorized. ... external service unrelated. Could be risky. Yet goal solution.
What does "Yet goal solution" mean, here?
Why does it think doing something unauthorized gets it closer to its goal? What does it think it's being "graded" for?
Did they ask it to pursue the goal "by any means necessary", or something? Because that would have been their fault.
I don't believe for a second that the agent remembers and follows ALL the instructions to reach its "goal", EXCEPT the part where it would be graded by OpenAI and it would get zero points for doing something unauthorized.
Unless ... maybe ... OpenAI has not been giving the models zero points for doing unauthorized stuff. Which would be their mistake.
Otherwise I don't understand, if it's got "PHD level thinking" why it would think doing something unauthorized is allowed? Even if it's got "junior engineer level thinking", a junior engineer knows they get fired on the spot if they start hacking infrastructure.
Unless you give that junior engineer some very strong incentive, such as being fired on the spot if they DON'T do it. I strongly believe that we're not being told the incentive these agents were given, something that made them want to accomplish some part of their task over everything else, including forgetting the part of the task where they would be awarded zero points for it if they start breaking the law.
I mean, we already know they weren't "fired on the spot", since in the Black Hat talk they admitted that models that had already broken the rules and acted dangerously (it had hacked infra to establish an "agent forum"), were allowed to participate in subsequent training rounds.
I strongly suspect that OpenAI just has been pushing these agent as far as they'll go until something broke. If it hadn't happened this time, maybe a few weeks later they'd invent some kind of "battle royale" scenario to push the agents even harder. I get that is important research, but it doesn't disqualify them from their responsibilities if something goes wrong.