4 ms·
> Monte Carlo Tree Search -- AlphaGo (although some of it is Neural Network training, the MCTS is the core of the algorithm) AFAIK, this is not correct. Many o
by bytefactory 9y ago
> Monte Carlo Tree Search -- AlphaGo (although some of it is Neural Network training, the MCTS is the core of the algorithm)
AFAIK, this is not correct. Many of the Go playing algorithms before AlphaGo used MCTS or some variant. The true breakthrough of AlphaGo was deep reinforcement learning.
> AlphaGo's performance without search
The AlphaGo team then tested the performance of the policy networks. At each move, they chose the actions that were predicted by the policy networks to give the highest likelihood of a win. Using this strategy, each move took only 3 ms to compute. They tested their best-performing policy network against Pachi, the strongest open-source Go program, and which relies on 100,000 simulations of MCTS at each turn. AlphaGo's policy network won 85% of the games against Pachi! I find this result truly remarkable. A fast feed-forward architecture (a convolutional network) was able to outperform a system that relies extensively on search.
https://www.tastehit.com/blog/google-deepmind-alphago-how-it-works/ https://www.tastehit.com/blog/google-deepmind-alphago-how-it...
I don't know whether AlphaGo Master (the next version of AlphaGo that was trained purely with self-played games and has not been beaten in 60+ games) even uses MTCS.