4 ms·
> A3C is known to not be suitable for adversarial environments Interesting! What are the main papers in this area? Any intuition why this is the case? is it b
by MasterScrat 6y ago
> A3C is known to not be suitable for adversarial environments
Interesting! What are the main papers in this area?
Any intuition why this is the case? is it because A2C generally results in brittle policies?
- fnbr 6y agoIt’s mostly an issue that A2C isn’t designed for adversarial environments. It also doesn’t have any notion of hidden information, while other algorithms (eg CFR) explicitly handle this. There’s a well-known phenomena of cycling, where agent A will beat agent B which beats agent C which beats agent A; A2C can exhibit this. Think of rock/paper/scissors- AlwaysRock beats AlwaysScissors which beats AlwaysPaper. To avoid this, you typically need to do some sort of averaging. The alphastar paper and blog post do a good job discussing these issues as they had similar problems. I’d say that’s a great starting point (and then following their references). Blog post: https://deepmind.com/blog/article/alphastar-mastering-real-time-strategy-game-starcraft-ii https://deepmind.com/blog/article/alphastar-mastering-real-t...