3 ms·
The first comment of that thread predicted AlphaGo! :O
by zump 10y ago
The first comment of that thread predicted AlphaGo! :O
- gwern 10y agoWell, so does this paper. Many people, myself included, thought the next step was a CNN+MCTS, because it was the obvious next step. It was doing it well and getting it all the way to superhuman quality (along with reinforcement learning/self-play) that made AlphaGo so important.
- dontreact 10y agoBiggest fundamental difference: no policy network learned via reinforcement learning. AlphaGo = MCTS + Value network + (Policy network). I think that that piece is pretty important and is what allowed AlphaGo to improve so much with self-play.
- rocqua 10y agoAparantly, just the policy network of alpha go plays at a decent level (around 1kyu - 1Dan IIRC). Alphago started with the policy network and later build the rest around that.