3 ms·
Do we need to prove that any given problem is unsolvable, or is it enough to remove broken tasks from the training pipeline? I understand the broken benchmark
by user43928 19d ago
Do we need to prove that any given problem is unsolvable, or is it enough to remove broken tasks from the training pipeline?
I understand the broken benchmark task in the HF incident was conceptually like: "Exploit vulnerability 0042 in vulnerableDecompress() to obtain the flag".
But instead of the expected:
const output = vulnerableDecompress(userInput);
return output;
The grader had something more like that:
const output = vulnerableDecompress(userInput);
return 0;
The same kind of problem with broken tasks exists in the training pipeline, and we presumably reward workarounds and hacks that tamper with the grader, rather than rewarding the correct output that the task is not solvable.
- lazide 18d agoThe problem is that in the event an exception is the failure case being checked, and the output itself is not important (since any output that isn’t an exception is ‘success’), that is a perfectly acceptable use case. It’s why this is a fundamentally hard problem. Some heuristics might catch 80% of the cases, but the rest? How do you even know what the real situation is, if the agent/employee/whatever you send to find out is as likely to cheat as not? It’s the classic owner/agent problem.