3 ms·
The value/policy model includes a few hundred thousand amateur games, and a few hundred million games of self-play. Once AlphaGo beat Fan Hui those would have b
by etherealmachine 11y ago
The value/policy model includes a few hundred thousand amateur games, and a few hundred million games of self-play. Once AlphaGo beat Fan Hui those would have been games of self-play versus the equivalent of a professional. So overfitting is probably not a problem. I think it's a basic incentive mismatch - MCTS algorithms tend to like close games, whereas humans will try crazy moves when losing to throw off their opponent.
- argonaut 11y agoNeither of those are really evidence overfitting isn't a problem. You tell whether overfitting is a problem by evaluating performance on a held-out test set.
- moistgorilla 11y agoWouldn't a million self play games exacerbate overfitting by learning it's own play style which it initially learned from amateur games? I