16 ms·
Our paper: http://www.nature.com/nature/journal/v529/n7587/full/nature16961.html http://www.nature.com/nature/journal/v529/n7587/full/nature1... Video from Nat
by Inufu 11y ago
Our paper: http://www.nature.com/nature/journal/v529/n7587/full/nature16961.html http://www.nature.com/nature/journal/v529/n7587/full/nature1...
Video from Nature: https://www.youtube.com/watch?v=g-dKXOlsf98&feature=youtu.be https://www.youtube.com/watch?v=g-dKXOlsf98&feature=youtu.be
Video from us at DeepMind: https://www.youtube.com/watch?v=SUbqykXVx0A https://www.youtube.com/watch?v=SUbqykXVx0A
edit: For those saying it's still a long way to beat the strongest player - we are playing Lee Sedol, probably the strongest Go player, in March: http://deepmind.com/alpha-go.html http://deepmind.com/alpha-go.html.
That site also has a link to the paper, scroll down to "Read about AlphaGo here".
If you want to view the sgfs in a browser, they are in my blog: http://www.furidamu.org/blog/2016/01/26/mastering-the-game-of-go-with-deep-neural-networks-and-tree-search/ http://www.furidamu.org/blog/2016/01/26/mastering-the-game-o...
- radicality 11y agoCan I read the paper somewhere without a Nature subscription?
- jimfleming 11y agoDeepMind/Google has a hosted copy: https://storage.googleapis.com/deepmind-data/assets/papers/deepmind-mastering-go.pdf https://storage.googleapis.com/deepmind-data/assets/papers/d...
- deleted 11y ago[deleted]
- deleted 11y ago[deleted]
- fitzwatermellow 11y agoOfficial Google Blog post also has some background info: AlphaGo: using machine learning to master the ancient game of Go https://googleblog.blogspot.com/2016/01/alphago-machine-learning-game-go.html https://googleblog.blogspot.com/2016/01/alphago-machine-lear... Interesting that DeepMind was using Google Cloud for compute. I imagine that the MCTS expansion can become massive. Any chance DeepMind may publish some of the internals about how many instances were used, how computation was distributed, any packages or frameworks used, etc. And congrats on achieving this impressive AI milestone!
- cfcef 11y ago> Any chance DeepMind may publish some of the internals about how many instances were used, how computation was distributed The computation requirements for training and running are given in the paper in both the body and the appendix details.
- fitzwatermellow 11y agoThanks, cfcef! Implementation details are in the "Methods" section. Have started experimenting with small GPU ML cloud jobs and the costs do add up. Wanted to get a sense what a large job looked like and indeed, AlphaGo is gargantuan. 50 GPUs approx train time one month for the policy/value network. So, a Google R&D size budget would be a prerequisite ;)
- tetraodonpuffer 11y agowas it easy to convince the players to have a match with AlphaGo? or was there some reluctance especially now when losing is becoming more of a possibility at even strength?
- eru 11y agoI don't know about Go professionals, but this might be the last time a human can win against computers. (Or the first time computers will win all the time.) It's a strange honour in either case.
- ewanmcteagle 11y agoWhen can we hope to see some form that will allow the public to play against this even if it is pay to play for each game or a weaker PC version? I hope the system does not end up being put away like Deep Blue was. Also, what kind of hardware results in this level of play and how is hardware correlated with strength here?
- rpearl 11y agoDeep Blue was only innovative in that it was specialized hardware for this type of search. The algorithms it used were well-established, and as there was no way to play it as a piece of hardware without great expense, there wasn't really a reason to keep it around. Chess engines you can run today, for free, on your own laptop, are far and away better than Deep Blue (and any human), and I believe still don't reach Deep Blue's raw speed.
- karussell 11y agoI'm curious: is there a Chess league for software? And if yes, how far are they already better (in ELOs) than humans if run on commodity server hardware?
- tetraodonpuffer 11y agoI think that would be https://icga.leidenuniv.nl/ https://icga.leidenuniv.nl/ I can't find the claimed ELO for Jonny (current champ) but Junior (previous champ) is listed at 3200+, Magnus' top rating, the highest ELO rating ever, is 2882 for reference
- bluecalm 11y agoThis is not a serious competiton. For serious stuff, see for example: 1)http://tcec.chessdom.com/ http://tcec.chessdom.com/ 2)http://www.computerchess.org.uk/ccrl/4040/ http://www.computerchess.org.uk/ccrl/4040/ There is a lot of politics in chess programming but the bottom line is that Komodo is currently the strongest program followed by Stockfish (which is distributed under GPL).
- sawwit 11y agoGreat achievement. To summarize, I believe what they do is roughly this: First, they take a large collection of Go moves from expert players and learn a mapping from position to moves (a policy) using a convolutional neural network that simply takes the 19 x 19 board as input. Then they refine a copy of this mapping using reinforcement learning by letting the program play against other instances of the same program: For that they additionally train a mapping from the position to a probability of how how likely it will result in winning the game (the value of that state). With these two networks they navigate through state-space: First they produce a couple of learned expert moves given the current state of the board with the first neural network. Then they check the values of these moves and branch out over the best ones (among other heuristics). When some termination criterion is met, they pick the first move of the best branch and then it's the other player's turn.
- sillysaurus3 11y agothey also train a mapping from the board state to a probability of how how likely it is a particular move will result in winning the game (the value of a particular move). How is this calculated? When some termination criterion is met Were these criterion learned automatically, or coded/tweaked manually?
- sawwit 11y ago1. The value network is trained with gradient descent to minimize the difference between predicted outcome of a certain board position and the final outcome of the game. Actually they use the refined policy network for this training; but the original policy turns out to perform better during simulation (they conjecture it is because it contains more creative moves which are kind of averaged out in the refined one). I'm wondering why the value network can be better trained with the refined policy network. 2. They just run a certain number of simulations, i.e. they compute n different branches all the way to the end of the game with various heuristics.
- someotheridiot 11y agoIf their learning material is based on expert human games, how can it ever get better than that?
- Florin_Andrei 11y agoWow, this is stunning. You guys beat a professional 2-dan player. That happened a lot sooner than expected. There's some kind of exponential evolution going on with AI these days. Is AlphaGo being made available to the public? I'm a mediocre player, but I'd like to try a few games against it. Current synthetic players don't quite play the same as humans do, and it's a bit jarring. I wonder if the Google AI is more human-like in its style. Anyway, here's the news from the p.o.v. of AGA: http://www.usgo.org/news/2016/01/alphago-beats-pro-5-0-in-major-ai-advance/ http://www.usgo.org/news/2016/01/alphago-beats-pro-5-0-in-ma...
- dorianm 11y agoDirect link to the paper: https://storage.googleapis.com/deepmind-data/assets/papers/deepmind-mastering-go.pdf https://storage.googleapis.com/deepmind-data/assets/papers/d...
- cfcef 11y agoWhen in March is the match? Will it be broadcast live?
- Inufu 11y agoThere will be a live broadcast on our YouTube channel, I think the date will be announced soon.
- mjmaher 11y agoWhat is your YouTube channel?
- fastturtle 11y agohttps://www.youtube.com/channel/UCP7jMXSY2xbc3KCAE0MHQ-A https://www.youtube.com/channel/UCP7jMXSY2xbc3KCAE0MHQ-A
- jermaink 11y agoShannon numbers: Chess: ~ 10^123. Go 19x19: ~ 10^360. Source: https://en.wikipedia.org/wiki/Game_complexity https://en.wikipedia.org/wiki/Game_complexity As a non-expert, may I ask (as the term does not appear in the paper): How valuable is the Shannon number in order to evaluate "complexity" in your context?
- shas3 11y agoSince both numbers are out of the realm of brute-forcing, the bigger achievement is because of the more fluid and strategic nature of Go compared to chess. Chess is more rigid than Go, and playing Go employs more 'human' intelligence than chess. Quoting from the OP paper: "During the match against Fan Hui, AlphaGo evaluated thousands of times fewer positions than Deep Blue did in its chess match against Kasparov; compensating by selecting those positions more intelligently, using the policy network, and evaluating them more precisely, using the value network—an approach that is perhaps closer to how humans play. Furthermore, while Deep Blue relied on a handcrafted evaluation function, the neural networks of AlphaGo are trained directly from gameplay purely through general-purpose supervised and reinforcement learning methods." "Go is exemplary in many ways of the difficulties faced by artificial intelligence: a challenging decision-making task, an intractable search space, and an optimal solution so complex it appears infeasible to directly approximate using a policy or value function. The previous major breakthrough in computer Go, the introduction of MCTS, led to corresponding advances in many other domains; for example, general game-playing, classical planning, partially observed planning, scheduling, and constraint satisfaction. By combining tree search with policy and value networks, AlphaGo has finally reached a professional level in Go, providing hope that human-level performance can now be achieved in other seemingly intractable artificial intelligence domains."
- mikekchar 11y agoI will admit to not following AI at all for about 20 years, so perhaps this is old hat now, but having separate policy networks and value networks is quite ingenious. I wonder how successful this would be at natural language generation. It reminds me of Krashen's theories of language acquisition where there is a "monitor" that gives you fuzzy matches on whether your sentences are correct or not. One of these days I'll have to read their paper.
- tzs 11y agoThis is a great advancement, and will be even more so if you can beat Lee Sedol. There are two interesting areas in computer game playing that I have not seem much research on. I'm curious if your group or anyone you know have looked into either of these. 1. How to play well at a level below full strength. In chess, for instance, it is no fun for most humans to play Stockfish or Komodo (the two strongest chess programs), because those programs will completely stomp them. It is most fun for a human to play someone around their own level. Most chess programs intended for playing (as opposed to just for analysis) let you set what level you want to play, but the results feel unnatural. What I mean by unnatural is that when a human who is around, say, a 1800 USCF rating asks a program to play at such a level, what typically happens is that the program plays most moves as if it were a super-GM, with a few terrible moves tossed in. A real 1800 will be more steady. He won't make any super-GM moves, but also won't make a lot of terrible moves. 2. How to explain to humans why a move is good. When I have Stockfish analyze a chess game, it can tell me that one move is better than another, but I often cannot figure out why. For instance, suppose it says that I should trade a knight for an enemy bishop. A human GM who tells me that is the right move would be able to tell me that it is because that leaves me with the two bishops, and tell me that because of the pawn structure we have in this specific game they will be very strong, and knights not so much. The GM could put it all in terms of various strategic considerations and goals, and from him I could learn things that let me figure out in the future such moves on my own. All Stockfish tells me is that it looked freaking far into the future and it makes any other move the opponent can force it into lines that won't come out as well. That gives me no insight into why that is better. With a lot of experimenting, trying out alternative lines and ideas against Stockfish, you can sometimes tease out what the strategic considerations and features of the position are that make it so that is the right move.
- Cyph0n 11y agoThe first point is a good one. As for the second point, I think that will be achieved only when we're close to realizing a general AI.
- jegutman 11y agoThis is actually a problem I've given a decent amount of though on (although not necessarily reaching a good conclusion), but I think these problems are actually related and not impossible for this simple case. It comes to an issue of what parts of the analysis and at what depth did a best move come to vision? Was it bad when it was sorted for 8 ply but good at 16? Maybe that won't "tell" a person why a move was good, but it gives a lot of tools to help try to understand them (which can be exceedingly difficult right now if a line is not part of the principal variation, but ultimately affects the evaluation by an existing "refutation". But I think the other "difficulty" is that 1800 players play badly in lots of different ways, 2200s play badly in lots of different ways and even Grandmasters play badly in lots of different ways, but very strong chess engines play badly only in a few sometimes limited ways.
- fspeech 11y agoAre games against Fan played by the distributed or single machine version of AlphaGo?
- fspeech 11y agoFound the answer to my question in the paper. It was the distributed version, as one would expect.
- conanbatt 11y agoHey Inufu. I just replayed the games and have to say that the first game the bot shows very high quality plays. The next two games, it seems like Fan Hui did not perform as well as the first (as opposed to the computer being clearly better than him). Where the games played in a row? Regardless, I'm looking forward to the games with Lee Sedol. I studied in his school in Korea, and personally know how hard it is to get to that level. My assessment is that the bot from those games will NOT beat Lee Sedol. So train it hard for march :)
- Inufu 11y agoThe games were all played on separate days. As Fan Hui mentioned in the video, he changed his strategy after the first game to fight more, so that may explain why it seems his performance changed.
- igravious 11y agoI've just replayed them as well. You can see that in the first game the bot played really solidly without risk-taking and didn't want to lose even a few stones. You could say that it played very conservatively but solidly. It won by a tiny margin (2.5) so Fan Hui probably concluded that the bot would win every game by a similar small margin if he didn't change the style of play. I'm sure that in the first game Fan Hui was sounding the bot out for strengths and weaknesses, seeing if it knew all the tesujis and josekis and whatnot. So from then on you see Fan Hui trying to mix it up and play more aggressively and what is very interesting is that he got outplayed in every game, even to the point of losing a big group and resigning. So - if you play conservatively and tentatively and solidly it'll beat you by a sliver, if you try to out-think it it'll nail you. At least at the 2dan pro level. I'd be hesitant for calling a Lee Sedol victory ahead of time. We know that in chess Kasparov beat IBM's bot initially but then IBM tweaked and within a couple of years the bot was too strong. Even though go is much harder than chess I predict that if Google lose this time and if they don't lose by much they'll win the time after that.
- pmoriarty 11y agoKasparov claimed that when IBM's team won against him they cheated.
- z0r 11y agoI believe that Lee Sedol would sweep all the games if he were to play Fan Hui under the same conditions, which makes the wait until the official match tantalizing. It's fair to say that AlphaGo has mastered Go, but there is a very large difference between a professional who has moved to the west and professional players who are competing at the highest level in regular matches. It's fair to represent Fan Hui as a master of the game, but misleading to represent him as equivalent to currently competing professionals. It is great that we'll get to see a match up against a player who is unquestionably one of the best of all time.
- beefman 11y agoKe Jie is the strongest go player. Sedol is the 5th strongest. http://www.goratings.org http://www.goratings.org The difference is 100 Elo, which is substantial. Edit: Ke Jie has 700 Elo on Fan Hui. About the same as the gap between Magnus Carlsen and the strongest player at my local chess club.
- apetresc 11y agoIn the recent 5-game match between Jie and Sedol a few weeks ago, it was decided in Ke Jie's favour by less than a single point in the fifth game. It literally would have come out differently if they'd used a subtly different (commonly used) scoring ruleset. It's not at all clear who's stronger at their peak.
- beefman 11y agoNo, but it's clear who's stronger on average.
- Rylinks 11y agoWasn't the margin in the 5th game b+1.5?
- DavidSJ 11y agoKeep in mind that professionals are counting throughout the game, and are playing to win, not to win by a lot. So a 0.5 point victory may simply mean the victor was confident in their position and chose not to take unnecessary risks.
- hyperpape 11y agoI follow professional go. It's true that you play safely when you're ahead, which narrows the margin, but .5 is still considered tightly fought.
- iopq 11y agoThat's not true, because Ke Jie would not have played the dame points if Japanese scoring would be used.
- thomasahle 11y agoIn the game against Lee Sedol, what hardware will AlphaGo be running on? Will it be a cluster or something consumer level? Also, will AlphaGo have access to Lee's previous games? Will Lee have access to AlphaGo's games for preparation?
- nopinsight 11y agoJapanese 9-dan pros and former Japanese cup holders who played against CrazyStone beat it less than 80% of the time [1], while AlphaGo's win rate against it is 80%, according to a comment below by inufu, an engineer at Google DeepMind. If transivity applies then AlphaGo is likely stronger than the average of those former Japanese champions, including Norimoto Yoda, who is currently ranked at 187th (about 300 Elo rating below Lee Sedol and 300 above Fan Hui) [2]. There's a saying in Go circles that there is a substantial gap in playing style or intuition between pros and even top-level amateurs. Whether that is true or not, AlphaGo has definitely crossed the threshold to pro-level play in Go. By March 2016, Google DeepMind would have improved AlphaGo somewhat at least through self-playing and perhaps more processing power. The game with Lee Sedol will be an interesting one to watch! [1] http://www.computer-go.info/h-c/ http://www.computer-go.info/h-c/ [2] http://www.goratings.org/ http://www.goratings.org/
- dfan 11y agoJust to clarify your comment so people aren't confused by the apparently low winning percentages, those are all 4-stone handicap games. (It's still an apples-to-apples comparison.)
- sehugg 11y agoI think the 80% win rate against CrazyStone is for the single-machine version. The distributed version won 100% of the time against CrazyStone and 80% of the time against the single-machine version.
- devy 11y agoFYI, Fan Hui, is ONLY a 2nd Dan Go player. [1] [1]: https://fr.wikipedia.org/wiki/Fan_Hui https://fr.wikipedia.org/wiki/Fan_Hui
- cardmagic 11y ago2 dan pro is a much stronger strength than 2 dan amateur. The difference between a 2 dan pro and a 9 dan pro is usually just one stone handicap whereas the difference between a 2 dan amateur and a 9 dan amateur would be around 7 stones.
- fspeech 11y agoThis is an impressive achievement. However there are many subtleties involved when humans play against computers. I think only time can tell how big a breakthrough this really is. It is telling that AlphaGo only won 3:2 for the informal games. As a computer doesn't know the difference between formal and informal this seems to indicate that Alpha isn't truly above Fan Hui in strength. Also the formal games were played with fast game rules, which may be particularly advantageous to computers. Unlike chess go accumulates pieces on board throughout the game. Towards the end there are many stones on board and it is easy for human to err while the search space (possible moves) actually gets smaller for the computer and there is no doubt that computer has stronger book keeping capabilities. So to fairly evaluate human vs computer we may need new time rules different from human vs human games. The paper does not disclose whether the trained program displays true understanding of game rules. Humans don't just use pattern recognition, they also use logic to evaluate game state. While this could be addressed by the search part of the algorithm the paper doesn't appear to give any indication on whether this was studied. For example, the board position strictly speaking does not determine the game state due to ko rules (so the premise of the paper that there is a valuation function v(s), where s is the board position, that determines game outcome is incorrect). It would be particularly interesting to see how the algorithm fares when there are multiple kos going on at the same time. Also it would be interesting to see how well the algorithm understands long range phenomenons such as ladder and liveness. With a million dollar challenge in the plan it is understandable the Google team may not want to disclose weaknesses of the algorithm but in the long run we will get to know how robust it really is. From my experience playing against conv nets I would say if you treat your computer opponent as a human it would be like playing against the hive mind of a group of experts with infallible memory and it is not to your advantage. So one would be better off trying "cheat" moves that human experts do not use on each other and see how well the computer generalizes. Without search and with neural nets alone it is clear that computers do not generalize that well. So it would be interesting to see how well search and neural nets work together and if someone could find the algorithm's weak spots.
- omphalos 11y agoYeah, MCTS is traditionally terrible with ko. It would be interesting to learn how AlphaGo addresses or doesn't address this.
- cowpig 11y agoCan you please please stop publishing in Nature (and behind paywalls in general)? I don't attend an expensive university but I like to learn things.
- pas 11y agoSomeone else linked the paper directly: https://storage.googleapis.com/deepmind-data/assets/papers/deepmind-mastering-go.pdf https://storage.googleapis.com/deepmind-data/assets/papers/d...
- shpx 11y agoWhere'd you guys get the 30 million Go games, and is there any chance that the rest of us can get that data to train our own nets?
- SonOfLilit 11y ago30 million moves. I'm pretty sure you can download that many from Go game databases listed here: http://senseis.xmp.net/?GoDatabases http://senseis.xmp.net/?GoDatabases
- thomasahle 11y agoIn the paper you estimate that Distributed Alpha Go has ELO around 3200 (if I read the plot correctly.) According to goratings.org, Fan Hui is rated 2900 and Lee Sedol is rated 3515. Doesn't that mean you still have work to do before beating Lee Sedol?
- jibalt 11y agoOf course they do, as they have said.
- tuxguy 11y agoamazing stuff Julian & congrats !!! is there any way to access the nature paper without a paywall(for academic, non-commerical use) ?