4 ms·
I considered clarifying that, but it's a bit subtle. First of all, neural networks are typically trained from (millions of) self-play games. So you have the "e
by codeflo 5y ago
I considered clarifying that, but it's a bit subtle.
First of all, neural networks are typically trained from (millions of) self-play games. So you have the "equally strong opponent" assumption right there in the weights of the network.
Second, when actually playing, the variations that are explored with high probability in MCTS are (recursively) informed by the same neural networks that are applied at the root. This means that there's no extra reward for moves that lead to complications in lines that only a "stupid human" would play.
In effect, neither alpha-beta search nor Monte-Carlo tree search have any concept of "tricking" a weaker opponent, admittedly for very different reasons and more probalistically in the MCTS case.
That's what I meant with "the implicit assumption that they play against the best response of an equally strong opponent" -- yes it's a simplification, but directionally true I think.