3 ms·
If you read the AGZ paper closely, they actually use checkpoints during training. Specifically, during training they only perform updates to the "stable" set of
by psb217 9y ago
If you read the AGZ paper closely, they actually use checkpoints during training. Specifically, during training they only perform updates to the "stable" set of parameters when the current "learning" set of parameters produces a policy which beats the stable set at least 55% of the time. The current stable parameters are what they use for generating the self-play data which they use to update the current "learning" parameters. I believe this is only mentioned in the supplementary material...
- gwern 9y agoI did and I would point out that while they use checkpoints, the training curves indicate this is not necessary, and what I meant is that they do not use the usual self-play (and evolutionary) mechanism of checkpoints from throughout the training history which is necessary to combat catastrophic forgetting (and which apparently wasn't enough to stabilize Zero on its own as it is the single most obvious thing to do but Silver notes all the pre-Zero self-plays diverged until they finally came up with that of MCTS supervision). The checkpoint mechanism there appears no more necessary than checkpoints in training any NN - it's critical to avoid a random error or bug wasting weeks of time but does not affect the training dynamics in any important way.
- gwern 9y ago(And Anthony et al 2017 don't use checkpoints at all, noting that it slows things down a lot for no benefit in their Hex agent.)