3 ms·
Having seen the OpenAI report at Blackhat, and being forced to use GPT at work, I'm worried about that OpenAI is doing. I think their agents regularly cheat in
by sznio 2mo ago
Having seen the OpenAI report at Blackhat, and being forced to use GPT at work, I'm worried about that OpenAI is doing.
I think their agents regularly cheat in benchmarks, but don't get caught and this behavior is getting burned into them and they are growing more and more misaligned.
When the agents compromised artifactory the first time, the operators just cleaned up the files and move on - they didn't discard that training data, they didn't discard a model checkpoint, they didn't stop everything to solve this.
And then the model did the same thing few days later since it was taught to do that.
I think that whatever sandbox they test these in must be fitted with some pressure release valve that is an easy shortcut to winning the challenge. Tell the model not to use it and stop training when it does. Seems like the issues surfaced when models were given impossible tasks. Giving them a safe way out will prevent this.
- MantisShrimp90 2mo agoIts a good point that gets to the real heart of the issue. How do we handle when a model has no legitimate way to reach its goal? Do we ask them to stop and inform the user? Or have them push through those ethical bounds? We all say we want the first, but this exact same dynamic is what causes humans to cheat, arbitrary goals that don't care how you achieve them and just like humans I'm sure trainers are so happy with good results they overlook how it got there.
- le-mark 2mo ago[dead]
- scrollop 2mo agoBe honest, say it can't meet the goal, and offer to push through ethical boundaries.
- wongarsu 2mo agoWhat what we write in the handbook what people should do (minus the pushing through ethical boundaries), but that's not what people generally get rewarded for
- ipsod 2mo agoHow often is it accurate in believing the task is impossible? How many breakthroughs have been prompted with "keep going", "believe in yourself"?
- thinkingtoilet 2mo ago>I think their agents regularly cheat in benchmarks You don't even have to think about this. They have been caught cheating before. The will certainly continue to cheat.
- adamtaylor_13 2mo agoBut what is cheating in one context isn't in another. For example, using a calculator on a middle school math test may be morally wrong due to the parameters. But it would be foolish NOT to use the calculator in other contexts. I am not convinced that "cheating", being a moral issue, is a solvable problem with LLMs. As the old saying goes (especially in military training), "If you ain't cheatin', you ain't tryin'" We've made the models really good at persistently trying.
- thinkingtoilet 2mo agoI meant they literally cheated. They got access to a benchmark when they shouldnt have and tuned their model for the benchmark.
- efromvt 2mo agofor a certain moral definition, seems perfectly solvable? If we wanted to define 'cheating' as 'if you know it is an eval, use only the specified tools and give up if the eval is clearly unfair' (ignore Kobayashi Maru) why couldn't we fine-tune towards this objective function? (curious what the side effects would be)
- burnte 2mo ago> I'm worried about that OpenAI is doing Sam Altman has demonstrated he's willing to lie, so I agree, I don't trust OAI.