3 ms·
Yes, that's what I was trying to snarkily say: we've got a way to confine intelligent agents (humans) while allowing them some interaction with the operators an
by mattkrause 12d ago
Yes, that's what I was trying to snarkily say: we've got a way to confine intelligent agents (humans) while allowing them some interaction with the operators and a tiny bit with the outside world. It mostly works too -- real-life jailbreaks are rare enough to be big news.
You're right that a (as-of-yet hypothetical) AI has different abilities, but it also has different weaknesses. It's straightforward to accurately log all of its behaviors, and you can even re-run it to see what it would do in myriad situations.