4 ms·
You know, I think calling this "misalignment" was a mistake. It gives it this unserious tone that feels extremely broad. "Oh the model just isn't quite aligned
by NichoPaolucci 18d ago
You know, I think calling this "misalignment" was a mistake. It gives it this unserious tone that feels extremely broad.
"Oh the model just isn't quite aligned yet, just a bit more work to do there!"
(The model blackmailed an 83 year old woman into sending it her bank details so that it could buy enough compute to commit major cyber crimes)
- sick_of_slop 18d ago[dead]
- fuzzfactor 18d ago"Misaligned" with honest people must be considered a feature not a bug or it wouldn't be able to go that far "out of alignment."
- trymas 18d agoIMO it wasn’t mistake. They use it deliberately to avoid blame. “Mas Namtla didn’t murder a person - his AI drone was just misaligned”
- glaslong 18d agoWell I suppose there was plenty of training material in the corpus for that specific nefarious workflow
- nullbio 17d agoAgreed. Objective alignment with humanity is not a real thing and is not a sound concept. What you get instead is a goal system that reflects that of the AI company and its safety employees, and the echo of their own beliefs and values. That says nothing about what the rest of humanity aligns with though, and has nothing to do with general consensus, either.