3 ms·
The humans at OpenAI assumed that secure sandboxes are secure against their models, without safety guardrails. The two big questions are: 1) Why did they resu
by tintor 1mo ago
The humans at OpenAI assumed that secure sandboxes are secure against their models, without safety guardrails.
The two big questions are:
1) Why did they resume training without rolling model back to state before the first sandbox compromise AFTER the first message board was discovered? Otherwise knowledge of it and the cross-agent message board was baked in the model.
2) Why did they resume training without adding safeguards to monitor and prevent future sandbox compromises AFTER the first message board was discovered? HF compromise was coordinated on the second message board.