3 ms·
They mentioned what the training input to AlphaGo was in the paper; a database of a few hundred thousand games from dan ranked players on KGS. This means mostly
by lambda 11y ago
They mentioned what the training input to AlphaGo was in the paper; a database of a few hundred thousand games from dan ranked players on KGS. This means mostly amateur players, though there are a few professionals who play on KGS as well.
However, this only gets you so far, and the training set is fairly small compared to what you want to really train both the policy and value networks well. So then they had it play millions of games against different versions of itself, training both a new policy network and the value network based on that.
It's unlikely that Lee Sedol's games made much if any impact on AlphaGo's training. It was bootstrapped off of high-level amateur, and some casual pro, play, but from then on it just trained against itself.