3 ms·
Humans play for a large lead because they don't have enough memory/power to accurately estimate the value of their positions, so they play for a buffer -- Alpha
by gertef 10y ago
Humans play for a large lead because they don't have enough memory/power to accurately estimate the value of their positions, so they play for a buffer -- AlphaGo has higher confidence in its valuation, so it can play it closer -- ~85% confidence of winning by 5 stones (with room for error) vs 99% chance of winning by 2 stones
- hyperpape 10y agoI believe that accuracy/"self-confidence" is part of it. However, I think it's also the case that AlphaGo has a monte carlo tree search in addition to the neural net, so it sometimes plays more conservatively than it needs to because it overweights obscure possibilities ("defending here is not necessary, but by doing so, I prevent some number of playouts where I play a dumb move and lose, and I can still win even if I defend"). Humans do the same thing, playing conservatively in a situation where they're far enough ahead. The difference is that a human sometimes looks at a move and says "This move works, and gains points. There is no risk." For a bot using MCTS, everything is a probability.
- taejo 10y agoAnother factor is that many micro-endgame sequences simply have a "correct answer" that loses the least points. Any human who's played Go for more than a few months knows these sequences, and if they choose to answer a certain move, will prefer the answer which is "always correct" to another move which would also win the game. This naturally leads to the winning player preserving their margin even when they could throw it away, while the machine has no such bias, and will just as happily throw away the margin as preserve it.