4 ms·
~One training pass later…~
by The28thDuck 4y ago
~One training pass later…~
- gwern 4y agoWould probably do little good. They show that the attack goes through almost perfectly even if you give the target an entire wallclock hour per move to do tree search to try to defend against it (as opposed to the more typical seconds or minutes per move search budget): https://goattack.far.ai/pdfs/go_attack_paper.pdf#page=8 https://goattack.far.ai/pdfs/go_attack_paper.pdf#page=8 So it's a serious blindspot.
- dpaleka 4y agoMore search won't do good, but why wouldn't targeted training help? The way I see it is that the adversarial policy search discovers positions which are off-distribution for anything seen in the victim's self-play training. But training on that particular sort of adversarial states should help against the human player which has learned the strategy, just like training on patch adversarial examples in vision helps against the same type of patches. Of course if the adversarial policy is again allowed to find off-distribution states (by playing against the victim), it will certainly find ways to beat it, until the model is playing perfectly. (Emergent gradient obfuscation could also theoretically happen, but I don't know if it has been demonstrated to actually happen.)
- ablob 4y agoMore targeted training won't do good, but why wouldn't more search help? We've apparently entered the stage where the deciding factor between who wins, man or machine, is just an arms race.
- dpaleka 4y agoMore targeted training won't do good, but why wouldn't more search help? My understanding is that gwern above linked solid evidence in the paper for more search not being enough, as in, the model's evaluation NN is so way off target when searching, that realistic amounts of search don't help. Go seems to have many possible moves per position, so the search doesn't go very deep anyway. Feel free to correct me if I'm wrong, it might be that I misremembered how AlphaGo-style systems work.
- AdamGleave 4y agoOne of the authors of the paper here! David Wu (aka lightvector), the creator of the KataGo AI system that we target, is actually doing a training pass now initializing self-play games to start at adversarial positions found by our adversarial policy. It does seem to have improved things significantly: our adversary's win rate goes down from >99% to around 4% against the latest KataGo checkpoint. However, the real question is whether KataGo has started properly understanding the cyclic groups, or has just learned some heuristics that saves it most of the time. In other words, if we repeated our attack again against the latest checkpoint, could it learn to reliably repeat the 4% of winning cases? We're looking into doing our own adversarial training run to see if we can get to a point where the KataGo agent is robust. My personal suspicion given how difficult adversarial examples have been to eliminate in image classifiers and other ML systems is that although we'll be able to train particular vulnerabilities out of the system, there's still going to be a long tail of issues that can be automatically discovered and exploited.
- gwern 4y agoYeah, it seems unlikely that some patchup can fix, in generality, such a deep blindspot. It reminds me of the problems AlphaGo had with ladders: I'm not sure DM ever really solved the ladder blindness, it just got good enough that they became irrelevant in practice.