9 ms·
Computers do what you say, not what you mean. If I write a function and name it quickSort, that's no guarantee that the function is a correctly implemented sor
by astrofinch 11y ago
Computers do what you say, not what you mean. If I write a function and name it quickSort, that's no guarantee that the function is a correctly implemented sorting algorithm. If I write a function called beNiceToHumans, that's no guarantee that the function is a correct implementation of being nice to humans.
It's relatively easy to formally describe what it means for a list to be sorted, and prove that a particular algorithm always sorts a list correctly. But it's next to impossible to formally describe what it means to be nice to humans, and proving the correctness of an algorithm that did this is also extremely difficult.
These considerations start to look really important if we're talking about an AI that's (a) significantly smarter than humans and (b) has some degree of autonomy (can creatively work to achieve goals, can modify its own code, has access to the Internet). And as soon as the knowledge of how to achieve (a) is widely available, some idiot will inevitably try adding (b).
Note: Elon Musk and Sam Altman apparently think spreading (a) to everyone is a good way to mitigate the problem I describe. This doesn't make sense to me. You can read my objections in detail here: https://news.ycombinator.com/item?id=10721621 https://news.ycombinator.com/item?id=10721621 There's another critique of their approach here: http://slatestarcodex.com/2015/12/17/should-ai-be-open/ http://slatestarcodex.com/2015/12/17/should-ai-be-open/
If you're interested to learn more, here's a good essay series on the topic of AI: http://waitbutwhy.com/2015/01/artificial-intelligence-revolution-1.html http://waitbutwhy.com/2015/01/artificial-intelligence-revolu...
- mikeash 11y agoI like the Paperclip Maximizer thought experiment to illustrate this: https://wiki.lesswrong.com/wiki/Paperclip_maximizer https://wiki.lesswrong.com/wiki/Paperclip_maximizer Short version: imagine you own a paperclip factory and you install a superhuman AI and tell it to maximize the number of paperclips it produces. Given that goal, it will eventually attempt to convert all matter in the universe into paperclips. Since some of that matter consists of humans and the things humans care about, this will inevitably lead to conflict.
- Udik 11y agoThe funny thing is that this "computers do what you say, not what you mean" comes directly from their lack of intelligence. So it's kind of strange that we talk about the threats of superintelligence brought along by the fact that, fundamentally, a machine is stupid. Am I the only one to see a slight contradiction there?
- Strilanc 11y agoGoals are orthogonal to intelligence. The fact that the AI understands what you want won't motivate it to change what it's optimizing. It's not being dumb, it's being literal. You asked it to make lots of paperclips, tossing you into an incinerator as fuel slightly increases the expected number of paper clips in the universe, so into the incinerator you go. Your complaints that you didn't mean that many paperclips are too little, too late. It's a paperclip-maximizer, not a complaint-minimizer. Choosing the goal for a superintelligent AI a goal is like choosing your wish for a monkey's paw[1][2]. You come up with some clever idea, like "make me happy" or "find out what makes me happy, then do that", but the process of mechanizing that goal introduces some weird corner case strategy that horrifies you while doing really well on the stated objective (e.g. wire-heading you, or disassembling you to do a really thorough analysis before moving on to step 2). 1: https://en.wikipedia.org/wiki/The_Monkey's_Paw https://en.wikipedia.org/wiki/The_Monkey's_Paw 2: http://lesswrong.com/lw/ld/the_hidden_complexity_of_wishes/ http://lesswrong.com/lw/ld/the_hidden_complexity_of_wishes/
- snowwrestler 11y agoThis reads to me like begging the question, by assuming the existence of a "superintelligent AI" without addressing how a goal-optimizing machine becomes a superintelligent AI in the first place. The exercise of fearing future AIs seems like the South Park underpants gnomes: 1. Work on goal-optimizing machinery. 2. ?? 3. Fear superintelligent AI. Or maybe it's like the courtroom scene in A Few Good Men: > If you ordered that Santiago wasn't to be touched, -- and your orders are always followed, -- then why was Santiago in danger? If a paperclip AI is so dedicated to the order to produce paperclips, why wouldn't it be just as dedicated to any other order? Like "don't throw me in that incinerator!"
- Strilanc 11y ago> assuming the existence of a "superintelligent AI" without addressing how a goal-optimizing machine becomes a superintelligent AI I'm just talking about the fallout if one did exist, saw ways to achieve goals that you didn't foresee, and did exactly what you asked it to do. I have no idea how the progression from better-than-humans-in-specific-cases to significantly-better-than-humans-at-planning-and-executing-in-the-real-world will play out. It's not relevant to what I'm claiming. > why wouldn't it be just as dedicated to any other order? It would be just as dedicated to those other orders. The problem is that we don't know how to write the right ones. "Don't throw me into that incinerator" is straightforward, but there's a billion ways for the AI to do horrible things. (A super-optimizer does horrible things by default because maximizing a function usually involves pushing variables to extreme values.) Listing all the ways to be horrible is hopeless. You need to communicate the general concept of not creating a dystopia. Which is safely-wishing-on-monkey's-paw hard.
- snowwrestler 11y ago> Computers do what you say, not what you mean. If we're going to start with that, then it has to apply to the full set of reasoning. Not just that computers will fail to consider whether to be nice to humans, but also that computers must therefore be explicitly told how to be effective in every particular way. If this remains true, then computers will not be resilient--their effectiveness will decline sharply outside of explicitly defined parameters. This is not a vision of terrifying force. Intuitively we can understand this by thinking about employees. One does exactly what he is told, but only what he is told, and then comes back for more instructions. Another can be given a goal, and then goes off and finds his own ways to accomplish that goal. Which one is more effective? Which one is more likely to compete for his manager's job some day? Put shortly: a computer that doesn't understand human society will not be able to make a significant independent impact on human society.
- edanm 11y ago"Put shortly: a computer that doesn't understand human society will not be able to make a significant independent impact on human society." Just like early humans who didn't understand animal's societies didn't have any impact? You're equating two different things which aren't necessarily equal - intelligence (in the sense of being able to achieve goals) and "agreeableness" to humanity. We could have one without the other. To use your analogy, an employee that is great at being given a goal and achieving it without explicit instructions, but doesn't necessarily have the same wellfare in mind as their boss.
- snowwrestler 11y agoWhat orders were early humans following?
- astrofinch 11y agoThe point is that humans have been able to destroy animal ecosystems to fit their own various ends without an in-depth understanding of those ecosystems.