2 ms·
GP seems to imply that should our view be true, then human jails should be empty by definition, because all prisoners are intelligent beings and would've eventu
by TeMPOraL 5d ago
GP seems to imply that should our view be true, then human jails should be empty by definition, because all prisoners are intelligent beings and would've eventually talked their way of it.
But this doesn't account for the fact that modern incarceration has built-in safeguards and mitigations based on centuries of cases of people talking, bribing or forcing their way out of prison, as well as getting outside assistance in forms ranging from lawyers to raiding parties equipped for demolition works. There are now procedural and technological means to prevent such incidents for happening, applied proportionally to the degree of risk.
Meanwhile, with AI, we're still at the point where everyone is assuming they can just lock the agent in a sandbox and prompt nicely to not poke at it too hard, and things will be fine. There's no multi-layered structural and procedural safeguards, and there's no recognition for the fact that AI operates faster than humans, and that quite likely it'll be smarter at this than average "jailer".
- mattkrause 5d agoYes, that's what I was trying to snarkily say: we've got a way to confine intelligent agents (humans) while allowing them some interaction with the operators and a tiny bit with the outside world. It mostly works too -- real-life jailbreaks are rare enough to be big news. You're right that a (as-of-yet hypothetical) AI has different abilities, but it also has different weaknesses. It's straightforward to accurately log all of its behaviors, and you can even re-run it to see what it would do in myriad situations.