4 ms·
The core technique of AlphaGo is using tree search as a "policy improvement operator". Tree search doesn't work on most real-world tasks: the "game state" is t
by panic 9y ago
The core technique of AlphaGo is using tree search as a "policy improvement operator". Tree search doesn't work on most real-world tasks: the "game state" is too complex, there are too many choices, it's hard to predict the full effect of any choice you might make, and there often isn't even a "win" or "lose" state which would let you stop your self-play.
- habitue 9y agoThis version explicitly does not use tree search.
- panic 9y agoMCTS means "Monte-Carlo Tree Search". It's the core of the algorithm. The big difference is that it doesn't use rollouts, or random play: it chooses where to expand the tree based only on the neural network.
- cjbprime 9y agoNo, 'habitue is correct. This new blog post says that the new software no longer does game readouts and just uses the neural net.
- Tarq0n 9y agoThat's not what Monte Carlo Tree search is. The new version is still one neural network + MCTS. There's no way to store enough information to judge the efficiency of every possible move in a neural network, therefore a second algorithm to simulate outcomes is necessary.
- Twirrim 9y agoRead the white paper. MCTS is still involved, right the way through.
- ankeshanand 9y agoThe new version does use MCTS, you should read the paper again. :)
- AlexCoventry 9y agoIt does, during training.
- panic 9y agoTree search is also used during play. In the paper, they pit the pure neural net against other versions of the algorithm -- it ends up slightly worse than the version that played Fan Hui, at about 3000 ELO.
- AlexCoventry 9y agoOh, so it's just not using rollouts to estimate the board position? Thanks for the clarification.
- mrec 9y agoIt doesn't use rollouts at all: > AlphaGo Zero does not use “rollouts” - fast, random games used by other Go programs to predict which player will win from the current board position. Instead, it relies on its high quality neural networks to evaluate positions.
- dastbe 9y agoIf you read the paper, they do in fact still use monte-Carlo tree search. They just simplify their usage in conjunction with reducing the number of neural networks to 1