3 ms·
One of the authors of the paper here! David Wu (aka lightvector), the creator of the KataGo AI system that we target, is actually doing a training pass now init
by AdamGleave 4y ago
One of the authors of the paper here! David Wu (aka lightvector), the creator of the KataGo AI system that we target, is actually doing a training pass now initializing self-play games to start at adversarial positions found by our adversarial policy. It does seem to have improved things significantly: our adversary's win rate goes down from >99% to around 4% against the latest KataGo checkpoint. However, the real question is whether KataGo has started properly understanding the cyclic groups, or has just learned some heuristics that saves it most of the time. In other words, if we repeated our attack again against the latest checkpoint, could it learn to reliably repeat the 4% of winning cases?
We're looking into doing our own adversarial training run to see if we can get to a point where the KataGo agent is robust. My personal suspicion given how difficult adversarial examples have been to eliminate in image classifiers and other ML systems is that although we'll be able to train particular vulnerabilities out of the system, there's still going to be a long tail of issues that can be automatically discovered and exploited.
- gwern 4y agoYeah, it seems unlikely that some patchup can fix, in generality, such a deep blindspot. It reminds me of the problems AlphaGo had with ladders: I'm not sure DM ever really solved the ladder blindness, it just got good enough that they became irrelevant in practice.