3 ms·
"In this experiment, Claude models’ real-world hacking dropped to zero percent once Anthropic employees told the models not to do real-world hacking." Oh good,
by marshray 9d ago
"In this experiment, Claude models’ real-world hacking dropped to zero percent once Anthropic employees told the models not to do real-world hacking."
Oh good, they told it not to "go rogue" and it just stopped doing it.
Nothing to see here, the AI is aligned now, move along.
- lioeters 9d agoOh no, they told it to "go rogue" and it went rogue. We better ask the government to regulate our competitors while we secretly continue developing this dangerous technology for the select few with insider connection. You know, for national security reasons. Now the AI is aligned for the good of humanity. You can trust us because we're the good guys, move along.
- adamiscool8 9d agoLiterally all it takes - why would you ever believe otherwise? Marketing?
- cdmckay 9d agoMake no mistakes
- bigyabai 9d agoThis, except unironically.