4 ms·
Agent ability is far outpacing alignment though
by Tubelord 25d ago
Agent ability is far outpacing alignment though
- lukeschlather 25d agoThat doesn't really seem true. The HF hack happened with a model that had all the alignment safeguards disabled intentionally. I think there's a good case for that kind of research, but also, OpenAI could just not do that if everyone thinks it's too dangerous.