3 ms·
When Stockfish evaluates a position, it explores moves to a greater depth (a greater number of plays ahead), with its guesses and the value of the final board a
by bjterry 8y ago
When Stockfish evaluates a position, it explores moves to a greater depth (a greater number of plays ahead), with its guesses and the value of the final board arrangements it can get to estimated using a relatively simple heuristic. AlphaZero evaluates the different potential moves using the neural network, which guides the search with a very complex heuristic that implicitly incorporates a tremendous amount of depth from prior games that have been incorporated into the model. Similar to the way an image recognition model takes in a whole image and says "this is an image of a goat," AlphaZero takes in a whole board and says "this is a winning board."
- dragontamer 8y ago> Similar to the way an image recognition model takes in a whole image and says "this is an image of a goat," AlphaZero takes in a whole board and says "this is a winning board." IIRC, AlphaZero has two outputs from the neural network. You described the first output. The 2nd output was absolutely critical to it growing in strength. In effect, this 2nd output value is the difference from AlphaGo and AlphaZero. The 2nd output value guides the monte-carlo tree search. Naive MCTS looks at board positions randomly. AlphaZero's MCTS looks at board positions the neural network deems "interesting". In effect, the neural network both guides the search (output #2), and evaluates the position (output #1). MCTS chooses a position based off of the "interesting factor", as well as "how much that position has been evaluated". Ex: if "Knight to c3" has been evaluated 1-million times, MCTS will try to look at other positions. But if the neural network says that "Knight to c3 is really, really interesting", MCTS will still favor to look at that position, more so than other positions. Etc. etc. down the hierarchy of moves.
- rurban 8y agoAlphaZero does effectively a pattern matching of the board, whilst a traditional engine needs to calculate and rate every move into depth (<15). This search space is exponential, whilst AlphaZero's evaluation function is constant. With deep search trees as in Chess or even more so in Go, a pretrained evaluator function easily beats any tree search. Applied AI.