14 ms·
One thing that isn't made clear in this writeup is that Master plays in a very nonhuman style, as opposed to the version of AlphaGo that beat Lee Sedol, which m
by dfan 10y ago
One thing that isn't made clear in this writeup is that Master plays in a very nonhuman style, as opposed to the version of AlphaGo that beat Lee Sedol, which mostly played like a strong human except for a few surprising moves. My first guess when I saw Master's games was that it was a program like AlphaGo that had its policy network trained from scratch rather than being bootstrapped by being given the goal of imitating the moves of strong humans. I'm eager to find out whether that was the case or whether it just moved away from human-like play during a very long self-play phase.
It really is a bit scary to see. I would not have guessed that human strategy was that deficient. Computer chess programs still tend to play with human-like strategy (partially because humans have coded their evaluation functions!) but godlike tactics. Master is not really playing like a human at all.
- halflings 10y agoCan you give some examples of what you consider non-human? Really curious since I don't know much about common strategies in Go.
- dfan 10y agoHuman moves tend to fit a narrative and be explainable, although often for very concrete reasons. For example: "I am sketching out territory while attacking an opponent group." "I am making my group safe so that I will not have to worry about its life while I accomplish other strategic goals." "I am making a very strong group so that I can use it to make it harder for my opponent to accomplish anything." Master, on the other hand, will sometimes seemingly just plop stones down in the middle of the board in a way that is hard to assign a narrative to. One can list a bunch of ways that the stone might come in handy in various futures, but it's too "vague" a move to be played with confidence by a human professional go player. Of course, human strategies may change in response...
- jorblumesea 10y agoThis is usually called the metagame and evolves over time. My guess is that due to these unseen tactics the metagame will change and human playstyle will become more AI-like and adapt to the changing circumstances. The real question is whether people would be able to keep pace with rapidly changing AI playstyles?
- Florin_Andrei 10y ago> The real question is whether people would be able to keep pace with rapidly changing AI playstyles? If we're far enough from "perfect play", then the answer is no.
- rtkwe 10y agoDepends on if people can figure out how to evaluate the more random seeming 'AI style' moves. With the complexity of Go it's possible to only way to really tackle the complexity is to wrap groups of moves together in a local narrative.
- Florin_Andrei 10y agoWhen playing against a much stronger player, it's always very hard to figure out why they play tenuki (make a move somewhere else on the board, apparently uncorrelated to the current fight). When AI gets strong enough (and it seems like it has already), it will just tenuki everyone all the time, while winning. Sounds like exactly what's happening already. It's past the event horizon for human understanding.
- Diederich 10y ago> past the event horizon for human understanding I think this phrase is going to pop up more and more frequently.
- eutectic 10y agoThis thread is literally the only Google search result for this phrase (for me)...
- Diederich 10y agoI searched for it in incognito mode and saw the same. Pretty interesting; it's such a nice, catchy phrase.
- gohrt 10y ago"Outside the light cone" is more on point -- it's far enough away or moving fast enough that we'll never be able catch up.
- cLeEOGPw 10y ago> It's past the event horizon for human understanding. AlphaGo creators could "rewind" the whole program state to that move and inspect the tree search probabilities according to the board states it looks through to find a list of board states that generate a cumulative highest probability of wining by doing a move in that exact odd spot. My guess would be that while humans tend to put stones with a single, double, or sometimes triple "reason", or in AI words, "high probability of local effectiveness in upcoming several turns", AlphaGo, with his ability to see further, can see past the local effectiveness into more global effectiveness and higher probability of winning further down the road. In other words, those tenuki moves might actually be past the event horizon of human understanding, but only if inspected by looking at them and thinking ourselves. If we use AlphaGo itself, it should be possible to find out the reason for every single tenuki it will ever do.
- seanwilson 10y ago> Human moves tend to fit a narrative and be explainable, although often for very concrete reasons. I'm not a Go player, but could it not be that if the AI played its moved in a more human order you could see what was going on and assign a narrative to the moves, but the AI can see the order of the moves doesn't matter sometimes so it seems to play more vaguely/randomly to observers? For example, say the AI played a set of moves at the top of the board that you could give an attacking narrative to and then it plays a set of moves at the bottom of the board that you could give an defensive narrative to. If you intermixed the move order for both narratives you, it seems like the AI is playing in a nonhuman way when really its moves have a narrative but humans are too fixated on the move ordering.
- Florin_Andrei 10y agoYou call it narrative when you understand what's going on. You call it chaos when it's far above your head.
- skmurphy 10y agoGood insight, I think there are implications for the Cynefin framework model that offers four possible state spaces: simple, complicated, complex, and chaotic. This classification may have as much to do with our state of information (or depth of understanding) as an objectively "chaotic state."
- xanderstrike 10y agoMore likely, playing it in order commits to a certain approach too early, and makes that approach predictable. Starting from the middle leaves other options open. Given that it trains against a copy of itself, with equivalent predictive powers, it makes sense that it would pick moves that have lots of branching possibilities because that would increase its effectiveness against itself.
- derefr 10y agoI've heard a similar approach described in military strategy at all levels: rather than looking for a single dominant tactic, you try with each "move" to create so many potentially-viable future positionings at once that your opponent cannot predict you in order to effectively concentrate their effort. It'd be very scary to watch a "sibling" to AlphaGo play a 4X game.
- tel 10y agoAlphaGo often plays surprising moves. Typically high-level commentary describes the horizon of "surprising" here as playing wider, faster, and more influence-oriented than current professionals think is appropriate. AlphaGo is repeatedly praised for identifying moves which seem too unconcerned with the opponents threats and instead take more power on the board in a leisurely fashion. The first big surprising move it played was a "4th line shoulder hit" which was thought to give too much territory to the opponent (a "3rd line shoulder hit" is considered a very fine and popular move). AlphaGo showed that at the right time even the 4th line variant is good since is gave AlphaGo large influence at the right time and position. A lot of go comes down to timing. AlphaGo doesn't seem to play moves which are inhuman in that they just make zero sense as much as plays moves which are more daring than humans would like to try. Then, infallibly, it turns out that AlphaGo's daring move had the mark of being amazingly well-timed.
- jonknee 10y ago> I would not have guessed that human strategy was that deficient. With as much freedom as Go allows I think it would be surprising if humans had stumbled upon an optimal strategy (or Master for that matter, I'm sure there is still much to be improved!).
- thefalcon 10y agoThe fact that Go commentators talk in terms of local strategy and narrative and anything other than the end-game from the very beginning made me feel fairly confident that Go was not a game that humans would ever reach optimal strategy levels at.
- Florin_Andrei 10y agoYou only need to do the math and look at the huge exponent to figure out that this is indeed the case.
- cLeEOGPw 10y agoAnd in addition to that, winning condition is also extremely fuzzy. Looking at the Master games, I couldn't tell why he is wining at all, granted I don't know much about Go. So if it requires some experience just to recognize a winner, and as far as I know sometimes even professionals can't tell for sure who is wining, it's pretty safe to say Go game is just too complex of a game for humans to come to optimal strategies in any reasonable amount of time.
- bluetwo 10y agoI don't know much about Go but I thought a comment about AlphaGo was interesting that it played in a way to marginally beat the player, which was different that most masters played, which is to clearly beat the opponent by as wide a margin as possible. Is this accurate? Does MasterP also use this style? Are there humans that can play this way? (I'm asking you because you seem to know what you are talking about here.)
- gertef 10y agoHumans play for a large lead because they don't have enough memory/power to accurately estimate the value of their positions, so they play for a buffer -- AlphaGo has higher confidence in its valuation, so it can play it closer -- ~85% confidence of winning by 5 stones (with room for error) vs 99% chance of winning by 2 stones
- hyperpape 10y agoI believe that accuracy/"self-confidence" is part of it. However, I think it's also the case that AlphaGo has a monte carlo tree search in addition to the neural net, so it sometimes plays more conservatively than it needs to because it overweights obscure possibilities ("defending here is not necessary, but by doing so, I prevent some number of playouts where I play a dumb move and lose, and I can still win even if I defend"). Humans do the same thing, playing conservatively in a situation where they're far enough ahead. The difference is that a human sometimes looks at a move and says "This move works, and gains points. There is no risk." For a bot using MCTS, everything is a probability.
- taejo 10y agoAnother factor is that many micro-endgame sequences simply have a "correct answer" that loses the least points. Any human who's played Go for more than a few months knows these sequences, and if they choose to answer a certain move, will prefer the answer which is "always correct" to another move which would also win the game. This naturally leads to the winning player preserving their margin even when they could throw it away, while the machine has no such bias, and will just as happily throw away the margin as preserve it.
- jkkramer 10y agoI would call its style unorthodox rather than nonhuman. It still plays common josekis (standard opening sequences) but often chooses uncommon variations. Its mid game is full of startling moves backed by VERY good reading. There's definitely still discernible strategy that us mortals can learn from. If I recall correctly, the version that beat Lee Sedol was trained on amateur games plus self-play. My guess would be that this new version relies more heavily on pro games.
- conistonwater 10y ago> Its mid game is full of startling moves backed by VERY good reading. This is pretty similar to what chess engines do.
- rictic 10y agoThe version that beat Lee Sedol was trained on pro games.
- jkkramer 10y agoGot a source? I can only find references to training on amateur games (e.g. https://en.wikipedia.org/wiki/AlphaGo_versus_Lee_Sedol#AlphaGo https://en.wikipedia.org/wiki/AlphaGo_versus_Lee_Sedol#Alpha...)
- hyperpape 10y agoNo, but I'm pretty sure the parent post is right. Pro games were included. Edit: I'm not actually that sure. I'm asking around right now (with go players, not DeepMind people).
- dfan 10y agoTheir Nature paper says "We trained the policy network p_sigma to classify positions according to expert moves played in the KGS data set. This data set contains 29.4 million positions from 160,000 games played by KGS 6 to 9 dan human players; 35.4% of the games are handicap games." It is possible that they fed it some pro games after the Fan Hui games but before the Lee Sedol games, but that would be weird; at that point it was already learning from self-play rather than trying to match human moves. That said, I don't think that Master's better performance comes from being trained on pro games. The AlphaGo version that played Lee Sedol played much more like a human pro than Master does.
- furyofantares 10y agoSince AlphaGo's original training was to predict human moves, it should know how surprising a move is in addition to knowing how strong it is. My thought at the time was it could improve its game against humans by giving surprise value some weight -- when it can pick a slightly weaker but highly surprising move, pick the surprise.
- whiteandnerdy 10y agoIt definitely does know how surprising a move is - it generates 'surprisingness' numbers for both itself and its opponent based on how likely it is a human pro would have come up with the move. I like your idea of maximising the 'alienness' of play though probably not to the exclusion of playing the strongest moves.
- tempestn 10y agoYou're probably right, if they were optimizing for short-term win rate vs humans. The real goal is to advance the state of the art in AI though, in which case you wouldn't want to use 'cheats' like that to make the problem easier.
- stouset 10y agoAs computers are able to evaluate positions faster (and therefore deeper), the "godlike" tactics are dominating over human-style strategy. It used to be that computers played "computer-like" moves because they didn't understand the position. Now, they play computer-like moves because "understanding" the position isn't as important as just being able to see 25+ moves ahead. In a nutshell, positional play in chess is simply heuristics we humans use to be able to evaluate a position in lieu of being able to calculate deep non-forced lines. Computers do use this to an extent (as you point out, we coded their evaluation functions) but positional play matters less when you see all the outcomes of every possible tactic with 100% accuracy. So computers tend to play reasonably human-like in the openings, but by the time you reach the middle game they'll happily enter lines where their pawn structures are shattered, pieces appear superficially to have little coordination, and where their king safety appears compromised (all things humans rarely intentionally do), all because they've seen that it works out 25+ moves in advance.
- 1024core 10y ago> Now, they play computer-like moves because "understanding" the position isn't as important as just being able to see 25+ moves ahead. I don't think you can "see 25+ moves ahead" in Go. The branchout factor is just too big.
- dfan 10y agoWell, stouset is talking about chess, but in Go, Monte Carlo Tree Search plays out all the way to the end of the game; it just does so in a much less exhaustive way, for exactly the reason you mention.
- stouset 10y agoMy comment was in a subthread about computers playing chess in a more human-like manner.
- skj 10y agoOne of the nice things about Go is that branches rejoin quite a bit, mitigating some of the branching factor concerns.
- stouset 10y agoAs computers are able to evaluate positions faster (and therefore deeper), the "godlike" tactics are dominating over human-style strategy. It used to be that computers played "computer-like" moves because they didn't understand the position. Now, they play computer-like moves because "understanding" the position isn't as important as just being able to see 25+ moves ahead. In a nutshell, positional play in chess is simply heuristics we humans use to be able to evaluate a position in lieu of being able to calculate deep non-forced lines. Computers do use this to an extent (as you point out, we coded their evaluation functions) but positional play matters less when you see all the outcomes of every possible tactic with 100% accuracy. So computers tend to play reasonably human-like in the openings, but by the time you reach the middle game they'll happily enter lines where their pawn structures are shattered, pieces appear superficially to have little coordination, and where their king safety appears compromised (all things humans rarely intentionally do), all because they've seen that it works out 25+ moves in advance.
- karpathy 10y agoAnother interesting aspect is that AlphaGo gets a larger advantage over human players if it plays novel moves and forces rare situations that the human players are not familiar with. So if AlphaGo is allowed to self-play a long time and "drift away" in strategy space from human games that could help a lot. This is even more the case in time-limited matches where the humans are forced to play intuitively in out of sample states.
- deleted 10y ago[deleted]
- zeven7 10y ago> had its policy network trained from scratch rather than being bootstrapped by being given the goal of imitating the moves of strong humans I just watched a couple of games, and I can't agree. Master's fuseki looks a lot like human fuseki. It plays some original, unusual, confusing, non-human looking moves, yes. It also plays some moves at what would normally be considered the "wrong time" conventionally. But the fuseki still looks highly derived from human games to me -- human with some (sometimes major) tweaks. When AlphaGo is trained fully from first principles alone, I expect another Shin Fuseki -- a revolution in opening strategy. I would be surprised if Go Seigen, Kitani Minoru, et al had finally discovered optimal opening strategy at the beginning of the 20th century (developments since then have mostly been small refinings of what was started in the Shin Fuseki era, not revolutionary).
- ogrisel 10y agoMaybe it does not use the same opening style when it plays against a version of itself. It would be very interesting to have pro-players comment on published records of alphago self-play. Maybe alphago has discovered a new balance between black and white (that is a new optimal value of the komi) but when playing with the human defined of the value of the komi its optimal style is also different than what it would be otherwise.
- Radim 10y agoYou mean, like these three AlphaGo self-play games from September? ;) https://deepmind.com/research/alphago/alphago-games-english/ https://deepmind.com/research/alphago/alphago-games-english/ (analysis by Gu Li and Zhou Ruiyang, two top pros; standard komi)
- hyperpape 10y agoNote that these games look much more human than the ones dfan was describing. There are surprising ideas, but they are still much more normal.
- 10y ago
- inimino 10y agoProfessional go player Otake Hideo supposedly said he would ask for three stones if playing against God. That was before AlphaGo. My guess has always been that the real gap is much higher, and a perfect player could give 6 or even 9 stones to the world's top players. In perfect play there would be no such thing as joseki. I don't think it would be a game we would even recognize.
- codehotter 10y agoAs play gets closer to optimal it gets more and more difficult to play more efficiently than your opponent. To play so much more efficiently than your opponent to overcome a handicap of 6 stones strikes me as extremely unlikely at pro level. At amateur level, say a 2d vs a 2k player would perhaps have a 99.9% winrate. To make it an even game, would take 4 handicap stones. But at pro level, I think it's possible for player A to have a 99.9% winrate against player B, but on two handicap stones, for player B to be the favorite. Even an engine that is enormously successful against top human players would struggle at high handicaps vs them.
- inimino 10y agoIn fact I agree, and if a perfect player were available, pro players would quickly get better at taking handicap! What I really wanted to express is the amount of headroom available between top players and perfect play. When I say nine stones, I mean, whatever the win rate is between an idealized 9p and 8p player, there would be nine more such steps between the 9p player and god. Probably more. I don't necessarily mean that god would have even odds giving nine stones to top players, because that's a different game. However, I have seen enough games where strong amateurs are taken apart with shockingly high handicaps by top pros, especially in faster games, to wonder. I think we simply fail to imagine how strange perfect play would be. Even if a pro spent the rest of their life thinking about the next move, they are unlikely to find the one true best move. A player that always played that move would be so far ahead of anything we've seen that we just can't imagine how much better it would be. Imagine knowing at move 10 that the best move, given perfect play by both, leads to a 3.5 point win in 256 more moves, while the second-best move leads to a 2.5 point win after 310 moves. Just stating it this way shows the amount of headroom there is above human play.
- xapata 10y agoThe same thing has happened in poker. The computer plays moves that are "obviously bad" according to human heuristics -- frequent limp opening and donk(ey) betting -- yet the computer is able to incorporate those moves into its strategy successfully.