3 ms·
One of the authors here! Great to see some discussion in the paper. Your summary of computer Go vs human rule sets seems right to me. But I think there might be
by AdamGleave 4y ago
One of the authors here! Great to see some discussion in the paper. Your summary of computer Go vs human rule sets seems right to me. But I think there might be a slight misunderstanding. We had friendlyPassOk set to false for all of our evaluation except one game which was played not against our adversarial policy, but one of my co-authors Tony who was trying to mimick the adversarial policy.
We evaluated KataGo under Tromp-Taylor with "self-play optimizations" described in https://lightvector.github.io/KataGo/rules.html https://lightvector.github.io/KataGo/rules.html which basically involves removing stones that can be proved to be dead using Benson's algorithm. This was the same evaluation used in the KataGo paper, and KataGo trained using these rule sets. (KataGo was also trained with some other rules -- it was randomized during training so it transfers across rules, and KataGo gets the rules as input.)
You might find this discussion of our paper at https://www.reddit.com/r/MachineLearning/comments/yjryrd/comment/iuqm4ye/?utm_source=reddit&utm_medium=web2x&context=3 https://www.reddit.com/r/MachineLearning/comments/yjryrd/com... by the lead author of KataGo interesting. He wasn't that concerned about the rule set, primary concern was that we evaluate in a low-search regime, which is a fair critique. But he overall agrees with our conclusion that self-play just cannot be relied upon to produce robust policies sufficiently OOD.