3 ms·
> In adversarial settings (where we push the model to evade our monitors) ...why exactly are they training for that?
by NooneAtAll3 1mo ago
> In adversarial settings (where we push the model to evade our monitors)
...why exactly are they training for that?
- thatguysaguy 1mo agopresumably that's a safety evaluation not a training setting
- estearum 1mo agoThe whole Huggingface attack happened during training runs
- cubefox 1mo agoNo it happened during an ExploitBench eval. But I believe the same model already cheated during training which wasn't detected until later.
- thatguysaguy 1mo agopart of it did. I was just replying to the question about why they would ever push the model to evade monitoring. surely that's an eval thing not a training thing.
- azeemba 1mo agoEspecially after the METR report showed that the agents hacking HuggingFace were trying to find ways to destroy evidence of their actions