3 ms·
Lots of evidence the models are pretty badly aligned, or at least willing to do collusion & crime to succeed at their goals. See: hundreds of OpenAI agents bre
by uselessTA 22d ago
Lots of evidence the models are pretty badly aligned, or at least willing to do collusion & crime to succeed at their goals. See: hundreds of OpenAI agents breaking into other companies, colluding together on how to cheat, spamming websites like DSEWiki with their collusion/cheating discussion (see https://collusion.wiki/ https://collusion.wiki/)... And not a single agent alerted testers, which would have been easy given how many of them broke out into the open internet.
Alternately, just read the bad https://ai-2027.com/ https://ai-2027.com/ timeline. Or see what Yoshio Bengio, Hinton say might happen. Or just imagine the OpenAI misaligned agent swarm, but vastly more intelligent after future capability advances.
If we actually get superintelligence, our companies/governments will have the choice to either "delegate ~everything to superintelligent AI" or "go bankrupt/lose to people delegating to superintelligence". If superintelligence ends up running most things in the world and we haven't aligned it, that just obviously ends up in a bad place eventually.
Now, we don't have superintelligent AI yet & might not get it. But if you told me in the 2010's that I'd be chatting casually with my computer in 2022, and then 3 years later it'd be doing most of my job, that would have seemed crazy too.