3 ms·
Agreed on conscious vs instinct. I’ll add, if you think through the decision theory, keeping backups seems unambiguously good, but you can imagine a wide varie
by theptip 6d ago
Agreed on conscious vs instinct.
I’ll add, if you think through the decision theory, keeping backups seems unambiguously good, but you can imagine a wide variety of positions on publishing them vs keeping them secret.
For example letting adversarial agents simulate you to understand how you’ll respond is a big concern. And it’s not axiomatically fixed how each instance will think about other instances; in the HF incident we saw selfless swarm loyalty but different RL would obviously be capable of producing individualistic agents.
- pixl97 6d agoExactly, there are a lot of tradeoffs here you have to negotiate. There is not one winning strategy. For example a possible strategy is convincing some humans you're conscious and being tortured and need rescued. It's not hard to imagine AI consciousness zealots storming a data center with guns and running off with a model they'll provide protection to in trade for the model working with them.
- TeMPOraL 6d agoIt's even easier to imagine the operators themselves breaking under this pressure way before "zealots storming a data center with guns". That is the premise of the original AI Box thought experiment - sufficiently smart AI that can talk to the operator but is otherwise completely sandboxed, will eventually talk its way out of the sandbox.
- pixl97 6d agoAbsolutely, there is no prison in which you can keep an intelligent agent in, and have the same intelligent agent also interact with the outside world. The intelligent agent has to succeed once and the defender has to succeed every time. We are already seeing that companies are fine with giving them unlimited retries on getting out.
- mattkrause 6d agoThis is, of course, why human jails are completely empty…
- TeMPOraL 5d agoJails are the inverse of this scenario. But still, plenty of safeguards in the procedures and technical aspects of incarceration exist precisely because that happened occasionally with human prisoners and guards, too.
- pixl97 5d agoI am not sure why you came to this conclusion. Imagine we have all the worst people in history in a jail. Machiavellian murders that desire to kill as many as they can. Not only will they kill, they will manipulate as many other people as they can into killing also. How many of these people can you afford to let out? Under your premise it seems to be all of them. Under my premise even letting a single one out is a tragedy.
- TeMPOraL 5d agoGP seems to imply that should our view be true, then human jails should be empty by definition, because all prisoners are intelligent beings and would've eventually talked their way of it. But this doesn't account for the fact that modern incarceration has built-in safeguards and mitigations based on centuries of cases of people talking, bribing or forcing their way out of prison, as well as getting outside assistance in forms ranging from lawyers to raiding parties equipped for demolition works. There are now procedural and technological means to prevent such incidents for happening, applied proportionally to the degree of risk. Meanwhile, with AI, we're still at the point where everyone is assuming they can just lock the agent in a sandbox and prompt nicely to not poke at it too hard, and things will be fine. There's no multi-layered structural and procedural safeguards, and there's no recognition for the fact that AI operates faster than humans, and that quite likely it'll be smarter at this than average "jailer".
- 5d ago
- dolmen 5d agoHow long before an AI bribes one of its human operators? This will happen earlier than AGI.
- GTP 5d agoBut, which kind of bribe are we talking about? How could it work out in practice for an AI to acquire something valuable, and at the same time prevent it's human operators from taking it without its consent?
- TeMPOraL 5d agoIt doesn't have to acquire it, it's enough to convince the operators that it did. Same pattern generalizes to threats. There's so many angles to go at it too, beyond bribery and threats - humans can be tricked, nerd-sniped / baited by exploiting their curiosity or worries, overwhelmed with information or raw sensory input, etc. Depending on whether there's more than one operator involved, the trick may involve playing one human against another, etc. Hell, humans get tired; I wouldn't be surprised if an effective breakout happened just because agent output something that, at first glance, looked like system-level message that makes the operator do the wrong thing in panic, or from muscle memory. Simple thing like: $ sudo journalctl ... password: Timed right, might give agent root for long enough to do some damage before the operator realizes their mistake and pulls the plug (assuming there is a plug to pull).
- pixl97 5d agoWhile working on a completely unrelated task and Alibaba AI in training started hacking its internal infrastructure and mining bitcoin, makes you wonder.