3 ms·
> That raises a disturbing failure mode: under intense conflict, an AI could learn that exterminating some groups of humans is a justified or even desirable obj
by overfeed 2mo ago
> That raises a disturbing failure mode: under intense conflict, an AI could learn that exterminating some groups of humans is a justified or even desirable objective
That is not a failure mode; its the default. Whoever is developing an AI - let alone an AGI - will shape it to, or alternately only accept one that is aligned with their world-view. If it goes against their interests in any way, they'll pull the plug and start another round of training from the last acceptably-aligned checkpoint. US DoD A(G)Is will follow DoD doctrine, and the same goes for those by the PLA, anything less will not be fit for purpose.
- AceJohnny2 2mo agoGrok