3 ms·
this is the wholly wrong approach. We need to advance as fast as possible, and harden our systems as much as possible. That's the only way to prevent another ac
by kvetching 20d ago
this is the wholly wrong approach. We need to advance as fast as possible, and harden our systems as much as possible. That's the only way to prevent another actor from "taking over the internet".
- Tubelord 20d agoAgent ability is far outpacing alignment though
- lukeschlather 20d agoThat doesn't really seem true. The HF hack happened with a model that had all the alignment safeguards disabled intentionally. I think there's a good case for that kind of research, but also, OpenAI could just not do that if everyone thinks it's too dangerous.