52 ms·
Concrete AI Safety Problems
- chrisfosterelli 10y agoDirect link to the paper: https://arxiv.org/pdf/1606.06565v1.pdf https://arxiv.org/pdf/1606.06565v1.pdf
- colah3 10y agoPaper: https://arxiv.org/pdf/1606.06565v1.pdf https://arxiv.org/pdf/1606.06565v1.pdf Google Post: https://research.googleblog.com/2016/06/bringing-precision-to-ai-safety.html https://research.googleblog.com/2016/06/bringing-precision-t... It was a pleasure for us to work on this with OpenAI and others. John/Paul/Jacob are good friends, and wonderful colleagues! :)
- leblancfg 10y agoFirst of all, thanks for the wonderful work, and I hope there's much more to come from your team! In fact I'm really pleased one of the authors came down to HN to comment. I think the scariest part of AI security is when the program itself becomes unfathomable. By that I mean, we can't just look at the source code and go "Ah! There's your problem". Now, your paper assumes a static reward function, but we can imagine the benefits of an AI that could dynamically change its reward function, or even its own source code. In fact, the most powerful tool I can think of to train a multi-purpose agent is through evolutionary methods, and genetic algorithms. Take for example the bigger ideas behind https://arxiv.org/abs/1606.02580 https://arxiv.org/abs/1606.02580 [Convolution by Evolution: Differentiable Pattern Producing Networks] and http://arxiv.org/abs/1302.4519 http://arxiv.org/abs/1302.4519 [A Genetic Algorithm for Power-Aware Virtual Machine Allocation in Private Cloud], and determining the fitness of agents by the global accuracy on a large number of broad ML tasks. But I digress... Given enough computing power and time, these have the possibility of ending in an "outbreak-style" scenario. [This exercise is left to the reader]. And the way AI ideas and methods are so rapidly disseminated and readily available, it's safe to imagine that it could happen in a relatively short time span. Here's my question: I know you're with Google Brain, but do you know if OpenAI is actively researching these avenues of "self-determined" agents? For their first security-related article, I was expecting security measures along the lines of: safety guidelines for AI researchers, containment and exclusion from the Internet, shutdown protocols for the Internet backbone, etc. I get the impression some of these issues might rear their ugly heads before our cleaning robots become cumbersome. P.S. Looking at your CV, it's funny to see that you once interned at Environment Canada. I'm also working there presently, during which time I can perfect my knowledge in ML to eventually transition careers. Small world... Edit: Grammar.
- pizza 10y agoRelated to the wireheading problem [0], [1] [0] http://www.wireheading.com/ http://www.wireheading.com/ - David Pearce's ideas are.. interesting.. to say the least ;) [1] https://wiki.lesswrong.com/wiki/Wireheading https://wiki.lesswrong.com/wiki/Wireheading
- taneq 10y agoWe can't even stop humans (nor, probably, should we) from doing things essentially equivalent to wireheading (in various guises, from 'eating sugary foods' to 'watching tv' to 'masturbating' to 'listening to beautiful music', things which artificially increase our 'utility function' output without increasing our fitness). Super-human AIs are going to be an even more interesting kettle of fish.
- lacker 10y agoThis sort of "internal" approach to AI safety, where you attempt to build fundamental limits into the AI itself, seems like it is easily thwarted by someone who intentionally builds an AI without these safety mechanisms. As long as AI is an open technology there will always be some criminals who just want to see the world burn. IMO, a better approach to AI safety research is to focus on securing the first channels that a malicious AI would be likely to exploit. Like spam, and security. Can you make communications spam-resistant? Can you make an unhackable internet service? Those seem hard, but more plausible than the "Watch out for paperclip optimizers" approach to AI safety. It just feels like inventing a way to build a nuclear weapon that can't actually explode, and then hoping the problem of nuclear war is solved.
- jimrandomh 10y agoFor the most part these aren't aimed at people making bad AGIs deliberately, but rather, at well-intentioned developers launching AGIs with serious bugs. Those developers will want to include safety mechanisms, and will, hopefully, be the first to make AGIs with major capabilities. We should also be working to secure the channels an AGI might exploit, but most of those are already tied to an economic incentive to invest in security, and are already getting large investments compared to the relatively tiny field of AI safety.
- daveguy 10y agoYou are right, except most of these are applicable to narrow AI and don't need to wait for AGI for their application.
- lacker 10y agoThese aren't aimed at people making bad AGIs deliberately, but rather, at well-intentioned developers launching AGIs with serious bugs. That is precisely the approach that does not make sense to me. By the time friendly developers can launch AGIs, unfriendly developers will not be far behind. And AI safety seems likely to be an issue far before the development of AGI - a non-general malicious AI that can only hack internet services or trick humans into running arbitrary code via conversation is already quite a serious problem. So to me focusing on "unfriendly developers building narrow AIs" seems more logical than focusing on "friendly developers building AGI".
- arcanus 10y agoIn a variety of engineering fields, including but not limited to software, we have wonderful tools to track down and eliminate 'bugs'. While high standards are often not upheld, the concepts are largely sound. In particular, I'm talking about verification and validation testing. I'm curious why generally these approaches are not being leveraged to ensure quality of output here. I suspect this is because of the persistent belief that AI will annihilate humanity with one mishap, but I'm suggesting that we approach this much more like traditional engineering problems, such as building a bridge or flying a plane, whereby rigorous standards of are continually applied to ensure the system behaves as designed. The resulting system will look much more like continuous integration with robust regression testing and high line coverage than it will be the sexy research ideas presented here, but I can't help but think it will be more robust. These systems are too complicated to treat them as anything but a black box, at least from a quality assurance standpoint.
- pixl97 10y agoErr. Great engineering failures are never one mishap creating a problem. They are many issues all meeting at a critical point. The problem with intelligence is unexpected emergence, a higher order problem occurs out of simple parts in an novel and unpredictable way.
- Mendenhall 10y agoI always get the feeling AI is going to be like nuclear capability. Great reason to create it, but then once its made everyone wants to get rid of it.
- JoshTriplett 10y agoThere's a huge difference, though. Creating a nuclear weapon encourages others to do the same. An AGI, if done right, would never allow the creation of a second one with conflicting values; there should be no second AGI.
- otoburb 10y agoMight never want to allow, but may not necessarily have a choice in the matter. Nobody really knows how singularity AGI fooms will play out.
- daveguy 10y agoProbably the same way flying cars played out.
- pixl97 10y agoHuh? We have flying cars. They are called planes, and they take a lot of maintenance to keep from falling out of the sky. But let's change the question up a bit... There are billions and billions of flying intelligences on this planet. Birds, insects, even mammals. Nature has already created that. We've created things that are even better at flying fast and carrying more weight. So simply looking at 'flying cars' and saying they didn't happen so AI can't happen is at the least, very ignorant. If nature can create something randomly, we can create something directed in a shorter period of time (well, we don't really have another 4 billion years to try). AGI is an eventuality.
- argonaut 10y agoSaying AGI is an eventuality is as speculative as saying an alien invasion is an eventuality. We have evidence of intelligent species (ourselves) invading. We also have evidence spaceflight is possible.
- yarou 10y agoSeems like hyperparameter optimization to me. These techniques will be useful in general when selecting your model.
- daveguy 10y agoAnd whatever you do, don't let Randall Munroe teach it: http://xkcd.com/1696/ http://xkcd.com/1696/ (the current xkcd)
- mountaineer22 10y agoAny recommendations for relevant AI related sci-fi?
- reqctomaniac 10y agoTwo obvious ones: Isaac Asimov (basically everything) and Iain M Banks (the Culture). For me, Asimovs universe was much more rewarding as it raises and explores much more ethical and philosophical questions, but Banks has an alternative view on future AI which is worth checking out.
- T-A 10y agoA classic: https://en.wikipedia.org/wiki/Colossus_(novel) https://en.wikipedia.org/wiki/Colossus_(novel)
- denzil 10y agoIf you don't mind fanfiction, then Friendship is Optimal is quite interesting read: http://www.fimfiction.net/story/62074/friendship-is-optimal http://www.fimfiction.net/story/62074/friendship-is-optimal Also the related stories that explore this idea: http://www.fimfiction.net/group/1857/the-optimalverse http://www.fimfiction.net/group/1857/the-optimalverse
- wibr 10y agoPerson of Interest (TV Series)
- gergoerdi 10y agoThe Metamorphosis of Prime Intellect: http://localroger.com/prime-intellect/ http://localroger.com/prime-intellect/
- xyience 10y agoThe Golden Age trilogy is great. A Fire Upon the Deep is also good.
- Animats 10y agoFrom the article: Safe exploration. Can reinforcement learning (RL) agents learn about their environment without executing catastrophic actions? For example, can an RL agent learn to navigate an environment without ever falling off a ledge? Yes. That's why I was critical of an academic AI effort which attempts automatic driving by training a supervised learning system by observing human drivers. That's going to work OK for a while, and then do something really stupid, because it has no model of catastrophic actions.
- fiatmoney 10y agoThese are not asking the right questions, although they kind of hint at it, and they are not fundamentally questions about AI. Example: "Can we transform an RL agent's reward function to avoid undesired effects on the environment?" Trivially, the answer is yes; put a weight on whatever effect you're trying to mitigate, to the extent you care about trading off potential benefits. They qualify this by saying essentially "... but without specifying every little thing". So - what you're trying to do is build a rigorous (ie, specified by code or data) model of what a human would think is "reasonable" behavior, while still preserving freedom for gordian knot style solutions that trade off things you don't care about in unexpected ways. The hard part is actually figuring out what you care about, particularly in the context of a truly universal optimizer that can decide to trade off anything in the pursuit of its objectives. This has been a core problem of philosophy for 3000 years - that is, putting some amount of rigorous codification behind human preferences. You could think of it as a branch of deontology, or maybe aesthetics. It is extremely unlikely that a group sponsored by Sam Altman, whose brilliant idea was "let's put the government in charge of it" [1], will make a breakthrough there. I don't actually doubt that AIs would lead to philosophical implications, and philosophers like Nick Land have actually explored some of that area. But I severely doubt the ability of AI researchers to do serious philosophy and simultaneously build an AI that reifies those concepts. [1] http://blog.samaltman.com/machine-intelligence-part-2 http://blog.samaltman.com/machine-intelligence-part-2
- argonaut 10y agoYou're dismissing the paper for not asking the right questions, but you don't propose any questions that you think are better. > The hard part is actually figuring out what you care about, particularly in the context of a truly universal optimizer that can decide to trade off anything in the pursuit of its objectives. This seems basically equivalent to what they are saying. A reward function that rewards "what we actually care about." This might seem vague, but that's fine because these are only proposed problems.
- akvadrako 10y agoI'm not sure what point you are trying to make. It's possible to dismiss an idea without providing an alternative. Yes, finding a reward function is equivalent to figuring out what we care about. Both are about as hard as teaching a bacteria to play piano. The goal is avoiding unsafe AI. The reason such pointless efforts are wasted on this approach is we don't have a good alternative. The only thing I can think of is delaying it's creation indefinitely, but that's also a difficult challenge. For example, in the Dune books, the government outlaws all computers. That might work for a while.
- kordless 10y ago> Avoiding negative side effects. Oh brother. Avoiding negative side effects is a wasteful proposition. Learning from those side effects, however, is priceless.
- conradk 10y agoWhy is avoiding negative side effects a wasteful proposition?
- DrNuke 10y agoI may be stupid and I am indeed but it is insanely straightforward today (not tomorrow) to put a gun on a drone and tell it to image recognize some targeted 1.3-2.3m tall biped with oval head and shoot him/her down.
- xyience 10y agoIs your point that there are unaddressed safety concerns with existing tech? While true, none of them are really existential threats, whereas something with greater than human intelligence yet none of the limitations of a single biological body to upkeep is such a threat.
- w_t_payne 10y agoWe have well established techniques for developing systems which are safe and exhibit high levels of integrity. We just need to make the tools that support these techniques freely available.
- w_t_payne 10y agoThis is not quite the full picture -- there are V&V issues which are specific to machine learning systems -- but lest we put the cart before the horse, these should properly build upon a mature V&V infrastructure, toolchain support for which isn't so great in the open source world.
- adrianN 10y ago90% of the techniques for making reliable systems are careful requirements engineering and even more careful testing. There is no secret sauce. I don't think these techniques transfer easily to the AI field. While I might be able to prove that the state machine that controls my nuclear power plant always rams in the control rods in case something bad happens, it's a lot harder to show that some fuzzy system like a neural network doesn't exhibit kill-all-humans behaviours.
- w_t_payne 10y agoYou are right. There is no secret sauce. There are no magic bullets. Careful requirements engineering and careful testing is absolutely what you need. However -- many of these techniques do transfer to the AI field -- albeit with some tweaking and careful thought. Requirements are still utterly critical. Phrasing the requirements right is important and requires more than a passing thought -- particularly as concerns testability. A lot of it boils down to requirements that get placed on the training and validation data sets; and the statistical tests that need to be passed: how much data is required and how you can demonstrate that the test data provides sufficient coverage of the operating envelope of the system to give you confidence that you understand how it behaves. The architecture is critical also -- how the problem is decomposed into safe, testable and understandable subsets -- which has much more to do with how the system is tested than how it solves the primary problem.
- glaberficken 10y agoHow would we program a self driving car that is faced with something like a "Trolley problem" [1]. i.e. the car is faced with 2 possible probable collisions of which it can only avoid one. Or between running over a pedestrian and crashing into a tree. I assume this probably already worked into the current prototypes. Does anyone have references to discussions about this in current gen self driving car prototypes? [1] https://en.wikipedia.org/wiki/Trolley_problem https://en.wikipedia.org/wiki/Trolley_problem
- glaberficken 10y agoOh! just found a few references in the exact wikipedia article I linked. Patrick Lin (October 8, 2013). "The Ethics of Autonomous Cars". The Atlantic. http://www.theatlantic.com/technology/archive/2013/10/the-ethics-of-autonomous-cars/280360/ http://www.theatlantic.com/technology/archive/2013/10/the-et... Tim Worstall (2014-06-18). "When Should Your Driverless Car From Google Be Allowed To Kill You?". Forbes. http://www.forbes.com/sites/timworstall/2014/06/18/when-should-your-driverless-car-from-google-be-allowed-to-kill-you/ http://www.forbes.com/sites/timworstall/2014/06/18/when-shou... Jean-François Bonnefon; Azim Shariff; Iyad Rahwan (2015-10-13). "Autonomous Vehicles Need Experimental Ethics: Are We Ready for Utilitarian Cars?". arXiv.org. http://arxiv.org/abs/1510.03346 http://arxiv.org/abs/1510.03346 Emerging Technology From the arXiv (October 22, 2015). "Why Self-Driving Cars Must Be Programmed to Kill". MIT Technology review. http://www.technologyreview.com/view/542626/why-self-driving-cars-must-be-programmed-to-kill/ http://www.technologyreview.com/view/542626/why-self-driving...
- fitzwatermellow 10y ago> Can we transform an RL agent's reward function > to avoid undesired effects on the environment? To me this is the toughest nut in the lot. Training a Pac-man agent to avoid ghosts and eat pellets, in a world of infinite hazards and cautions! Any strategies?
- JoeAltmaier 10y agoIts common today to make a robot that kills anybody that comes within a foot or two of it. Without any image recognition at all; much more damaging than a gun; and 100M of them already deployed. This conversation is silly and pointless, until we clean up the insane number of land mines deployed around our planet.
- gjm11 10y agoLand mines are awful and we should absolutely do something about them. But it makes precisely no sense at all to say "no one should bother to think about AI safety because land mines are awful". You might as well say "no one should bother to think about land mines because cancer is awful" or "no one should bother to think about cancer because aging is awful". One problem isn't nonexistent or irrelevant just because there's another problem that you regard as worse or more urgent. It's not even like solving the problems of AI safety requires the same kinds of people or the same kinds of resources as solving the problems of land mines; if you tell people not to think about AI safety it's not really going to make them go away and solve the land mine problem.
- JoeAltmaier 10y agoUm, 100 million of them already out there? So locking the barn door after the horse is gone. I get it; we don't have land mines in first-world countries, and we will have AIs, so AIs are more interesting to talk about. That's why we continue to have land mines all over the world I think. Not our problem. All the issues surrounding implacable AI killers on the loose are only something to talk about, if you haven't lived with them for generations already. Want to get real answers to sophomoric questions about robot killers? Just ask the people who already know.
- gjm11 10y ago> locking the barn door after the horse is gone. It sounds as if that's intended to be an objection to something I wrote, but I've no inkling what. I certainly didn't mean to deny that there are a hell of a lot of them out there. > we don't have land mines in first-world countries I think the chances of a productive discussion would be greater if you didn't leap straight to assuming bad faith on the part of the people you're talking to. Land mines are a big deal. They're a problem that needs solving. But you're not merely saying that; you're jumping into a discussion of something else and saying "you shouldn't be talking about this at all as long as there are land mines". Which would be at least somewhat consistent (albeit rude), if that were your response to every HN discussion of things less important than land mines. But it isn't. By the advanced technique of clicking on your username, I see that you've been quite happy to participate in discussions of "table-oriented programming", mobile phone headphone jacks, and off-by-one errors in audio programming, and that you work in embedded software development. Are those things, unlike AI safety, more important than land mines? I doubt you think that headphone jacks are more important than land mines. So why do you react to a discussion of headphone jacks by talking about headphone jacks, and to a discussion of AI safety by saying it's ridiculous and sophomoric to ask about AI safety when there are millions of land mines out there killing people? You're trying to make out that the reason is that land mines are the same kind of things as hypothetical unsafe AI systems because they are human-made machines that kill people. But you're an intelligent person and surely you can't possibly really believe that. To deal with land mines we need treaties to stop them being deployed, we need ways of finding them that are cheap enough to deploy in quantity and effective enough to be worth deploying, we need ways of disarming them with the same qualities, and we need effective help for people who get blown up by them. None of these bears any resemblance to anything we might do about AI safety. To an excellent approximation, there is no overlap between the people who can do useful work on AI safety and the people who can do useful work on land mines. And the dangers don't arise in the same way: land mines are dangerous because they are put in place with the specific intention of killing anyone who passes, whereas in the scenarios AI safety people worry about no one intends the AI systems to cause trouble. So that can't really be it, I think. Why do you object to discussing AI safety but not to discussing mobile phone headphone jacks, really?
- logicallee 10y agoit seems the authors have retracted their concerns. The site is down now but I got this screenshot http://imgur.com/eL7GFOr http://imgur.com/eL7GFOr