3 ms·
Biggest fundamental difference: no policy network learned via reinforcement learning. AlphaGo = MCTS + Value network + (Policy network). I think that that piece
by dontreact 10y ago
Biggest fundamental difference: no policy network learned via reinforcement learning. AlphaGo = MCTS + Value network + (Policy network). I think that that piece is pretty important and is what allowed AlphaGo to improve so much with self-play.
- rocqua 10y agoAparantly, just the policy network of alpha go plays at a decent level (around 1kyu - 1Dan IIRC). Alphago started with the policy network and later build the rest around that.