7 ms·
There is a value network which estimates the win rate for the current player from a given board state, and a policy network which estimates the probability that
by eutectic 9y ago
There is a value network which estimates the win rate for the current player from a given board state, and a policy network which estimates the probability that each move should be played. As of the more recent iterations, these networks share their bottom layers for greater computational and training efficiency.
The value network is simply updated to match the real outcomes of games of self-play.
The policy network is updated to match the results of a tree search; for each board position many thousands of lines are explored using the value and policy networks, and then the policy is updated to match (a somewhat 'sharpened' version of) the number of lines in which each move was chosen.
When exploring each lines, at each step the move 'a' is chosen from the current board state 's' which maximizes Q(s, a) + P(s, a) / (1 + N(s, a)), where R is the policy network, Q is the average value network evaluation for lines where 'a' was picked from 's' (with the appropriate signs to match the current player at 's'), and N is the number of simulations in which 'a' was picked. When we reach an unseen board state, it is evaluated with the policy network and a new line is explored from the root.
This is less circular than it may seem because:
a) The value network is trained using real outcomes.
b) Towards the end of the game, the tree search sees real outcomes.
This training procedure allows the network to 'bootstrap', learning progressively more complex knowledge about how to play effectively.
- soVeryTired 9y agoI think the separation of value and policy networks was a feature of the older AlphaGo systems, but not AlphaGo Zero
- seanwilson 9y agoThanks. How does this play out when you're training against a single self played game then? Does it play a whole game with its current networks and then once it knows the winner it goes over each move after to train itself? > and a policy network which estimates the probability that each move should be played. So for this network, the input is the before and after board state and the output is the probability that this move should be played?
- eutectic 9y agoThe input is the current board-state and the output is the probability of each move.
- eutectic 9y agoOops, I obviously meant 'P is the policy'