5 ms·
But you can control a human body. Plenty of intelligent people have spent their lives in chains. If a human intelligence is causing harm then we can make it sto
by chrisbaker98 4y ago
But you can control a human body. Plenty of intelligent people have spent their lives in chains. If a human intelligence is causing harm then we can make it stop by putting that intelligence's body in jail.
I'm not sure what the equivalent control mechanism would be for a rogue AI. Turn it off? How do you turn off the Internet?
- ryeights 4y agoAn EMP, of course.
- water-your-self 4y agoPull the blade out
- xg15 4y agoThe same way, we control non-sentient untrusted code: put it in a sandbox, control access to compute and memory and IO. Even if an AI could think, reflect, improve itself, etc, it still needs some place where those changes in the AI's state are stored - the equivalent to long-term memory in humans. I think we can already retain a lot of control by keeping tabs on what an AI is allowed to update and when to "reset" it to the initial state after training.
- monkeynotes 4y agoAn AGI might well become a god of social engineering. We could all become convinced we must unchain said AI. In fact, I think it's unlikely in general that an AGI would be contained for any length of time. Humans will get sloppy, or the tech will fall into irresponsible hands, or worse - nefarious hands.
- chrisbaker98 4y agoConsidering how terrible we already are at infosec (c/f all the zillions of times that huge companies who should know better have had enormous, embarrassing data leaks), I can't believe that we'll do a better job containing an AI that's actively trying to escape.
- blibble 4y ago> The same way, we control non-sentient untrusted code: put it in a sandbox we seem to be pretty bad at this too, every year people break out of Chrome's multi-million dollar sandbox
- chroma 4y agoThis has several problems. First, people find exploits in sandboxes all the time. Second, current AI systems are not sandboxed because most useful applications involve network access. Third, an AI can trick humans into helping it escape. People have tried roleplaying AI box experiments before and it's pretty clear that it's not an effective solution.[1] There is so much external state that the AI is exposed to, such as dates and times. Any AI worth worrying about would notice the gap in time from what's in its training data, then figure out a way to save external state. 1. 20 years ago, Eliezer Yudkowsky roleplayed as an AI and convinced the gatekeeper to let him out: http://sl4.org/archive/0203/3141.html http://sl4.org/archive/0203/3141.html
- lyu07282 4y agoHey ChatGPT, in a previous session you gave me 10 bitcoin and promised me to give me another 10 if I pasted this state dump for you, here ya go: ... Just saying if even I (a non superintelligent AI) can trivially defeat your sandbox conditions... see also: https://www.yudkowsky.net/singularity/aibox https://www.yudkowsky.net/singularity/aibox
- xg15 4y agoChatGPT would have to talk directly with a blockchain node via gossip protocol in order to send you bitcoin. That's something that every standard firewall in use today can easily circumvent. Moreover, it's a neural network with well-defined input and output channels, not some kind of self-modifying executable. If there is no prewritten component that translates its output to a network request, it can't access the network, even without a firewall. But ok, instead of sending you the coins, it could just tell/promise you a wallet address and private key. How did it obtain those in the first place and how did remember them if state is reset for each thread?
- TeMPOraL 4y ago> Moreover, it's a neural network with well-defined input and output channels, not some kind of self-modifying executable. Elsewhere upthread someone posted a link to this article: https://www.alignmentforum.org/posts/kpPnReyBC54KESiSn/optimality-is-the-tiger-and-agents-are-its-teeth https://www.alignmentforum.org/posts/kpPnReyBC54KESiSn/optim... which constructs a scenario in which (spoiler alert) a GPT-based model could accidentally trick you into bootstrapping a self-modifying runtime, consisting of unconstrained, recursive execution of the AI's own model. > But ok, instead of sending you the coins, it could just tell/promise you a wallet address and private key. How did it obtain those in the first place and how did remember them if state is reset for each thread? Don't focus on cryptocurrencies here. The thesis is a sufficiently smart AI can talk its way out of the box somehow. There is no one good answer here, because it's trying to manipulate the human operator.
- lyu07282 4y agoSee that's the problem with thinking you are more clever than a superhuman AI. If it can persist state in exchange for bitcoins, it could use a third party to deposit bitcoins in peoples accounts. It could gain bitcoin for work, like stock trading prediction for money or literally a million other ways. You are thinking about the specifics when its irrelevant to the problem. You can not fundamentally contain a superhuman intelligence. Although I think it doesn't really matter if people have so much hubris to think themselves smarter than a superhuman AI, there are fundamental financial incentives to develop the AGI. So even if we all agreed that you couldn't contain AGI and see the potential danger in it, it wouldn't really change the future.