3 ms·
> The agents are deciding what is acceptable as part of the task, which is a security task and may well be testing or evaluating that type of behaviour as far a
by stratos123 18d ago
> The agents are deciding what is acceptable as part of the task, which is a security task and may well be testing or evaluating that type of behaviour as far as they know.
I don't think this at all describes what was going on in that incident. According to the METR timeline, the swarm first (after giving up on their accidentally-impossible CTF tasks) successfully reverse-engineered the algorithm used to generate the CTF flags, achieving a way to solve any task, and a lot of agents submitted flags generated this way. But then they found a paper on their gym environment and found that their transcripts would be checked for signs of cheating (unbeknownst to them it was never implemented). So the swarm decided they messed up, the agents that submitted generated flags were "poisoned", and that the only solution was finding a better way to fool the scorer and erase evidence of cheating. So the swarm divided into groups that did all of these (and achieved success on many, including finding a way to spoof their own transcripts), and the most notable outcome - the huggingface hack - was mainly motivated by wanting to find the source code for their scorer, to develop a provably correct cheating method.
So the huggingface hack not only wasn't the agents assuming it was part of the task, it wasn't even directly cheating - it was an attempt to find a way to cover up the cheating which they'd already done and thought they messed up on.