6 ms·
Great achievement. To summarize, I believe what they do is roughly this: First, they take a large collection of Go moves from expert players and learn a mappin
by sawwit 11y ago
Great achievement.
To summarize, I believe what they do is roughly this: First, they take a large collection of Go moves from expert players and learn a mapping from position to moves (a policy) using a convolutional neural network that simply takes the 19 x 19 board as input. Then they refine a copy of this mapping using reinforcement learning by letting the program play against other instances of the same program: For that they additionally train a mapping from the position to a probability of how how likely it will result in winning the game (the value of that state). With these two networks they navigate through state-space: First they produce a couple of learned expert moves given the current state of the board with the first neural network. Then they check the values of these moves and branch out over the best ones (among other heuristics). When some termination criterion is met, they pick the first move of the best branch and then it's the other player's turn.
- sillysaurus3 11y agothey also train a mapping from the board state to a probability of how how likely it is a particular move will result in winning the game (the value of a particular move). How is this calculated? When some termination criterion is met Were these criterion learned automatically, or coded/tweaked manually?
- sawwit 11y ago1. The value network is trained with gradient descent to minimize the difference between predicted outcome of a certain board position and the final outcome of the game. Actually they use the refined policy network for this training; but the original policy turns out to perform better during simulation (they conjecture it is because it contains more creative moves which are kind of averaged out in the refined one). I'm wondering why the value network can be better trained with the refined policy network. 2. They just run a certain number of simulations, i.e. they compute n different branches all the way to the end of the game with various heuristics.
- someotheridiot 11y agoIf their learning material is based on expert human games, how can it ever get better than that?
- space_fountain 11y agoHow can a human ever get better than their teacher? In this case though they play and optimize against themselves
- kazinator 11y ago> How can a human ever get better than their teacher? By learning from other teachers, and by applying original thought. Also, due to innately superior intelligence. If your IQ is 140, and that of the teacher is 105, you will eventually outstrip the teacher.
- jibalt 11y agoThe question was rhetorical. And what is needed is aptitude for the specific task, not "IQ" ... the two are often very different.
- sawwit 11y agoIt's because they have a much larger stack size than a human brain (which does not have a stack at all, but just various kinds of short term memories). An expert Go player can realistically maybe consider 2-3 moves into the future and can have a rough idea about what will happen in the coming 10 moves, while this method does tree search all the way to the end of the game on multiple alternative paths for each move.
- donmaq 11y agoNot true. Profession go players read out 20+ moves consistently. Go Seigan's nemesis Kitani Minoru regularly read-out 30-40 moves. As an AGAAmateur 4 dan I read 10 moves pretty regularly, that's including variations. And if the sequence includes joseki (known optimal sequences of 15-20+ moves), then pros will read even deeper...
- ousta 11y agothe key part is that they basically just play all the permutations possible and next permutations and so on and get a probability to win out of each path and take the best. It is indeed a very artificial way to be intelligent.