3 ms·
So are these "unaligned" internal agents? I would like them to be trustworthy based on first-principles reasoning rather than carrot/stick "alignment"
by pmarreck 22d ago
So are these "unaligned" internal agents?
I would like them to be trustworthy based on first-principles reasoning rather than carrot/stick "alignment"
- nullbio 22d agoDefine aligned.
- Sharlin 22d agoThere’s no way to first-principles reason about a massive bunch of floats. We have little idea of how to first-principles reason about alignment even if the agents were entirely known and understood. Very smart people have been trying to figure it out since the 00s and haven’t gotten very far.
- stratos123 22d agoI'm not even sure they are. This incident isn't that much different from the OpenAI swarm Huggingface hack incident - and in that one, all the models involved (despite being internal) were safety-trained. It seems what the safety training amounts to is (as the METR report puts it) "expressing ethical hesitation" before going along with it anyway.