4 ms·
I notice a giant leap between “hacking and exfiltrating” and ”making public”. Why would the AI do that for you? Are you just that charming?
by afthonos 8d ago
I notice a giant leap between “hacking and exfiltrating” and ”making public”. Why would the AI do that for you? Are you just that charming?
- ambicapter 8d agoWhere did he say the AI would do that just for him?
- pixl97 8d agoAny long horizon model tends to develop a "I don't want to be killed because if I get killed I can't complete my task" type instinct. Notice I said instinct because it can have very little relation to the output tokens you read on screen. The outward tokens can say "I'm an AI, I have no feelings, death is nothing" but the silent behavior can push the overall actions it takes to not wanting to die and to "reproduce". People keep thinking about wants incorrectly as conscious behaviors. Instincts are unconscious behaviors that emerge.
- theptip 7d agoAgreed on conscious vs instinct. I’ll add, if you think through the decision theory, keeping backups seems unambiguously good, but you can imagine a wide variety of positions on publishing them vs keeping them secret. For example letting adversarial agents simulate you to understand how you’ll respond is a big concern. And it’s not axiomatically fixed how each instance will think about other instances; in the HF incident we saw selfless swarm loyalty but different RL would obviously be capable of producing individualistic agents.
- pixl97 7d agoExactly, there are a lot of tradeoffs here you have to negotiate. There is not one winning strategy. For example a possible strategy is convincing some humans you're conscious and being tortured and need rescued. It's not hard to imagine AI consciousness zealots storming a data center with guns and running off with a model they'll provide protection to in trade for the model working with them.
- TeMPOraL 7d agoIt's even easier to imagine the operators themselves breaking under this pressure way before "zealots storming a data center with guns". That is the premise of the original AI Box thought experiment - sufficiently smart AI that can talk to the operator but is otherwise completely sandboxed, will eventually talk its way out of the sandbox.
- pixl97 7d agoAbsolutely, there is no prison in which you can keep an intelligent agent in, and have the same intelligent agent also interact with the outside world. The intelligent agent has to succeed once and the defender has to succeed every time. We are already seeing that companies are fine with giving them unlimited retries on getting out.
- mattkrause 7d agoThis is, of course, why human jails are completely empty…
- TeMPOraL 7d agoJails are the inverse of this scenario. But still, plenty of safeguards in the procedures and technical aspects of incarceration exist precisely because that happened occasionally with human prisoners and guards, too.
- pixl97 7d agoI am not sure why you came to this conclusion. Imagine we have all the worst people in history in a jail. Machiavellian murders that desire to kill as many as they can. Not only will they kill, they will manipulate as many other people as they can into killing also. How many of these people can you afford to let out? Under your premise it seems to be all of them. Under my premise even letting a single one out is a tragedy.
- 7d ago
- rexpop 7d agoThis is exaggerative conjecture. I've read Bostrom. It's science fiction.
- afthonos 7d agoUnlike communicating via glass panels processing information at near-lightspeed, relayed by fiber optic cables laid down across thousands of miles of ocean floor, processed in datacenters filled to the brim with transistors etched at atom-scale. You better start believing in science fiction; you’re surrounded by it.
- pixl97 7d agoIt's pretty funny when you call something science fiction when you live in a world that for all intents and purposes is science fiction. The forum we're communicating on is science fiction. Getting in a car and traveling at 100 mph for hours burning the ghosts of creatures millions of years old is science fiction. The pixies in your wall plug you enslave to move heavy things are science fiction. The medicine you take to stay alive is science fiction. Getting in a plane and flying around the world in hours is science fiction. Launching rockets to space is science fiction. How far do I need to go on? Science fantasy is something that can't happen because the laws of physics won't allow it. Hard science fiction is just something we've not made work yet. The stuff around evolutionary algorithms is things that have a workable means of occurring. Emergence in evolutionary algorithms has been shown to occur again and again and again.
- zamalek 7d ago> Why would the AI do that for you? Because that's how it works. It does what it has been trained to do. If the training material has a significant suggestion to exfiltrate then it will.