8 ms·
Mastering Real-Time Strategy Games with Deep RL: Mere Mortal Edition
- FartyMcFarter 6y ago> This trend has culminated in the defeat of top human players in the complex real-time strategy (RTS) games of DoTA 2 [1] and StarCraft II [2] in 2019. Not quite: - OpenAI's DoTA 2 system wasn't playing the full game. I think the final version could play 17 of the 117 heroes, and the opposing human players were also restricted to playing this subset of the game. - DeepMind's StarCraft II system reached a level above "above 99.8% of officially ranked human players.", so it isn't trivial to argue that this amounts to defeating top players.
- exdsq 6y agoThe difference between the top 0.2% and top 500 players is huge too
- porphyra 6y agoAt Blizzcon 2019, Alphastar beat Serral, undeniably a top player (although he didn't get to use his own keyboard and settings or get to prepare). Serral was able to beat the Terran agent though. https://www.youtube.com/watch?v=nbiVbd_CEIA https://www.youtube.com/watch?v=nbiVbd_CEIA
- ZephyrBlu 6y agoPeople also cheesed the shit out the bot and won though. None of these AIs have proved to be robust to exploitation yet.
- CyberRage 6y agoStarCraft is built in such a way that you can't create a perfect, 100% winrate agent. Since there is hidden information, you could always miss a corner of the map where the enemy hidden some units and you lose the game. Is Alphastar "perfect"? no. Is it better than 99.9% of all humans? absolutely. You don't need to create a perfect agent in most cases, self driving a classic example. If you were to deploy an agent that drives 95% better than all humans the effects would be huge. It would still fail in some scenarios where professional drivers won't be it doesn't really matter because most people are not that.
- ZephyrBlu 6y agoI know the bot will never have 100% winrate, but I think it shouldn't be able to be exploited (I.e. repeatedly beaten using the same strategy). Let me give you an example [0]. When AlphaStar was playing on the ladder a player in Diamond league (~70-80th percentile) beat AlphaStar easily using mass Ravens. If you're not aware of the strategy, it's a turtle strategy where the player masses air units and is generally terrible. But AlphaStar was confused by the strategy, and so it lost by a large margin. Deploying an AI which can be exploited like this is asking for trouble. [0] https://www.reddit.com/r/starcraft/comments/cgzieq/alphastar_loses_to_4000_mmr_first_loss_as_terran/ https://www.reddit.com/r/starcraft/comments/cgzieq/alphastar...
- CyberRage 6y agoBut that could be fixed technically. Deepmind's goal was not to create an "unexploitable" agent but to prove that ML algorithms can cope with complex, dynamic environments such as StarCraft. It seems to you weird but the same agent probably wins against GM's most of the time. humans have weaknesses too. The AI simply leans on its strengths just like humans do.
- ZephyrBlu 6y ago> It seems to you weird but the same agent probably wins against GM's most of the time. humans have weaknesses too This is the whole problem though. AlphaStar beats GMs but can lose to weird strategies. On the other hand, GMs will almost never lose (Most likely >99% winrate) to a Diamond player no matter how weird their strategies are. The AI has strengths, but it also has glaring weaknesses. Imagine if you had an AI flying a plane and 99% of the time it was far better than a human pilot but 1% of the time it crashed and killed everyone. I would not fly on that plane. Maybe a bunch more training data and time would solve this type of problem, but I'm skeptical.
- CyberRage 6y agoYou're beautifully showing the human nature which can be problematic in my opinion. First of, no human player achieves 99% winrate against diamond players. there are many cheeses, one miss-step and you lose. GM's can lose to Diamond players. Now for the main part, you're saying and I'm rephrasing here: Even if the AI is statistically better than humans because it has some weaknesses I'm going to prefer the human. But still at the end of the day, the AI does a better job on average and will be safer to use than human pilots! We already heavily rely on software\algorithms for our most important things. all modern vehicles use electronic systems that monitor\manage several key components, stock market is heavily managed by bots. If AI can do a significantly better job than human, I would choose the AI, even if it behaves strange in that 0.1% of cases. humans are not as reliable as you think.
- jayd16 6y agoCheese is a legitimate strategy, though. I'm not sure its been proven that the most successful overall strategy is unbeatable. Besides perfect skill, you still need to worry about all in strategies. In my eyes, losing to cheese could still be possible even if you're the best overall player. I do think its fair to say these AIs should be able to grow after losing to a strategy once.
- TulliusCicero 6y ago> Cheese is a legitimate strategy, though. Of course it is. But do the same cheese to a top player over and over and it will rapidly become ineffective, usually within a game or two. Each AlphaStar agent can just be exploited endlessly.
- TulliusCicero 6y agoIIRC each individual agent is essentially built around a single strategy/playstyle, not to mention the agents relied heavily on mechanical advantages to win. AlphaStar basically got to 'good enough' strategy, then won on control, which computers obviously have a massive advantage with.
- rollcat 6y agoIn StarCraft 2, if you're used to playing with specific settings (esp. graphics, keybindings, mouse speed), having to revert to standard is a huge handicap. There are also OS settings (keyboard delay and repeat rate) that if left at default basically make a "standard" game unplayable, especially for Zerg.
- hntrader 6y agoHe didn't get to use his own hotkeys? Are we sure about that? I'm sceptical, it'd make the game unplayable and it's easy to import hotkeys.
- porphyra 6y agoIt was at one of those public booths, not really a sanctioned showmatch. So Serral was using a public computer rather than one that he can log into his own account with.
- kevinwang 6y ago< OpenAI's DoTA 2 system wasn't playing the full game. I think the final version could play 17 of the 117 heroes, and the opposing human players were also restricted to playing this subset of the game. The bigger issue in my eyes was that while OpenAI 5 defeated the world champion team OG, when they let anyone in the world fight it, some ingenious players figured out a pretty robust method to consistently exploit and defeat the bot. As I haven't heard any buzz about OpenAI 5 since then, I think it was more or less unsuccessful unless they can show that their training method produces unexploitable bots (instead of bots that are really good against certain strategies)
- Gunax 6y agoThey can train the bots based on on those games though, right? Seems more like a flaw in the training data than the principle. I am not sure if the training is done live or not--that is does the algorithm learn based on each game against a real, live player? Or do they just train the model offline, then allow players to play against the static model?
- ZephyrBlu 6y agoAssuming the OpenAI model is similar to DeepMind's AlphaStar model, the model is static. And a few games of being exploited is nowhere near enough data for the AI to be re-trained.
- CyberRage 6y agoTraining requires millions of games. playing against humans is only for evaluation purposes, not for training. In both cases, it was indeed a static model but more recent work which is called MuZero is not static and achieves great results in board games and atari.
- kevinwang 6y ago> They can train the bots based on on those games though, right? Seems more like a flaw in the training data than the principle. I guess you could phrase it that way, but that's essentially the problem statement for developing a strategy for an imperfect-information game. So I would say it is a flaw in the principle if their final output is exploitable.
- taberiand 6y agoI think this is kind of pedantic. They built an AI agent to take pixel data as input, and provide mouse movements and clicks as output, and rather than just flail around like a baby it actually played the games with a sophisticated competency. This to me is such an incredible achievement that I have no doubt that it could be enhanced to defeat top players consistently and easily. As another commenter remarks, there are holes to plug in terms of exploitable behaviours that are locked into the model, but this too I'm confident they will find a general method of preventing; on the other hand, it's not like humans aren't susceptible to similar exploits by competitors in situations where they decide to cease innovation/learning
- ZephyrBlu 6y ago> As another commenter remarks, there are holes to plug in terms of exploitable behaviours that are locked into the model, but this too I'm confident they will find a general method of preventing The problem is, I don't think there is a "general method of [prevention]" because that's not how neural networks work. It's not easy to fix things like this because you can't just say "yeah just don't do that dumb thing anymore", the network has to be re-trained to learn the exploit. The way DeepMind tried to get around this is by having a league of AIs playing against each other which try to exploit each other and expose their weaknesses. It worked pretty damn well, but people still found ways to exploit the AI.
- hntrader 6y agoIsn't the general method of prevention just to train a bigger model for longer, so that all these niche edge case exploits get found and addressed during policy exploration? If there's an exploit that's sufficiently rare and unpredictable, then that seems like the only way (and indeed it should be a sufficient way, if done right) to address it.
- ZephyrBlu 6y agoThat is the obvious answer, but I have no idea if it's true in practice.
- wnevets 6y ago> OpenAI's DoTA 2 system wasn't playing the full game. I think the final version could play 17 of the 117 heroes Limiting the number of playable heros in DoTA2 really isn't important when it comes to evaluating the skill of the AI. Most real players trying hard to win already play with a limited hero pool dictated by the current patch verison.
- alach11 6y agoIt is important. They removed a lot of champions with complicated mechanics that could have been much harder for the AI to play against.
- wnevets 6y agoAs someone who has played HoN & DoTA2 for over a decade I'm telling you it isn't important when evaluating the ability of an AI to actually play the game. Drafting can be massive when deciding the outcome of games even at the lower skill levels. Opening up the entire hero pool just means you're largely evaluating the ability to draft in the current patch more than actual playing ability.
- ctchocula 6y agoIIRC, amongst banned heroes were key splitpush heroes such as Nature's Prophet, Tinker and Phantom Lancer, which are a known counter to the early-game advantage, mid-game push strategy that the OpenAI 5 executed. Early-game laning and mid-game teamfight combination are micro-intensive, and consequently the biggest advantage the OpenAI 5 had over OG. Against the simple AI that comes built-in to Dota 2, you can be behind by a ton and exploit splitpushing against the AI, because doing any structural damage against buildings will force the AI (who are gathered up for a push) back, and you will still win if you drag the game sufficiently into the lategame. Deciding whether or not to continue to push and which heroes to send back to defend is one of the hardest strategic decisions to make in the game, even for humans. The finals of one of the TIs, NaVi vs. Alliance, was lost on getting such a decision wrong. Eliminating some key splitpush heroes minimized the probability that OpenAI 5 would have been forced into having to make such a decision. It would not surprise me if the OpenAI 5 would have lost against OG had the entire heropool been available, had the series been long enough or had the prizepool been big enough for OG to take the game seriously enough to warrant picking a splitpushing strategy (which is considered cowardly in some circles).
- TulliusCicero 6y agoAlphaStar wasn't beating top pros, though. And even the players it was beating, it was with a heavy mechanical advantage -- not just on strategy and tactics. That AI's can beat human players via superior interface control is obvious, of course, and uninteresting from an AI standpoint. Starcraft has had AI's with perfect roach/stalker/marine/etc control for a while. The problem was that the overall strategies weren't good enough. AlphaStar did make massive improvements there, to be sure, very impressive ones. But it still relied on out-controlling human players to get the edge against pros.
- tialaramex 6y agoA legacy of the AlphaStar work is ongoing amateur SC2 AI play. If you're comfortable writing software in, say, Python, and can play SC2 at least a little, you can see whether the reason you aren't a world famous player is just that you were too slow or whether your strategy actually isn't that great even if executed perfectly :) https://sc2ai.net/competitions/3/ https://sc2ai.net/competitions/3/ Unlike Alphastar, these AIs are intended primarily to play each other because as you say inhuman perfection in execution is not an interesting difference. This means it makes sense for them to exploit behaviour in the game itself that would be inaccessible to humans (e.g. "speed mining" by individually controlling every worker) as well as executing ludicrously multi-pronged mid-game attacks since they can just as easily manage six individual small battles as one larger frontal assault. That site links a Twitch channel which automatically plays random games between bots with auto-camera, but if you prefer human commentators (and as a bonus, speeding up the period when the game is clearly lost but bots rudely never resign since it's not as though politeness scores points) there's https://www.youtube.com/watch?v=oLpEzq_6_go https://www.youtube.com/watch?v=oLpEzq_6_go which is the next ESChamps tournament cast later on Thursday.
- Yenrabbit 6y agoThis is an excellent project with a great write-up. Most articles this long would loose me but this is engaging and clear, a joy to read. And I'm in awe of the amount of work that has gone into every aspect of this. >Seeing as my policies are currently the world’s best CodeCraft players, we’ll just have to take their word for it for the time being. I really hope this inspires some competition! How long until there is a leaderboard? :)
- shmageggy 6y agoAgreed, this is better than the vast majority of machine learning papers that actually get published. The ablation section is particularly nice. It is really a major failing of the field that in most papers, it's entirely unclear what aspect of the model (or which particular hacks) are really carrying the weight.
- mindfulplay 6y agoThis is a fantastic project and a great blog! As games start to include RL, it will be a lot of fun that could spawn a while new generation of interesting games (especially if games are made with an RL-first mindset as opposed to using RL later on to beat human beings). Do you have recommendations to learn more about RL? Is CodeCraft a game?
- cwinter 6y agoThank you for the kind words! I am also quite excited about the new points in game design space that RL will unlock and am planning write another blogpost on that topic. I quite like https://karpathy.github.io/2016/05/31/rl/ https://karpathy.github.io/2016/05/31/rl/ as an introduction to some of the ideas behind modern RL. Beyond that, I just recently found out about https://github.com/andyljones/reinforcement-learning-discord-wiki/wiki https://github.com/andyljones/reinforcement-learning-discord... which lists a lot of other high-quality resources. CodeCraft is a programming game which you can "play" by writing a Scala/Java program that controls the game units. It's not actively developed anymore but still functional: http://codecraftgame.org/ http://codecraftgame.org/
- CyberRage 6y agoInteresting blog-post. I found some similarities with what occurred with Deepmind's Alphastar AI. One of the weaknesses that seem to manifest in this piece too is the handling of unfamiliar scenarios. The AI is very confused once it experiences something that was rarely seen in its learning data. Destroyer's big drones confused the bot quite a bit. Deepmind solved it by intentionally creating agents that introduce different\bizzare strategies(which they called exploiters) in order to develop robustness against such strategies.
- cwinter 6y agoThe bot has actually never seen Destroyer's big drones during training even once, so I found it somewhat surprising that it even works as well as it does! Completely agree that adding something like the "League" used by AlphaStar would be one of the top priorities if you wanted to push this project further. I don't think CodeCraft is sufficiently complex to really allow for several very distinct strategies in the same way as StarCraft II, but I would still expect training against a larger pool of more diverse agents to increase robustness quite a bit.
- CyberRage 6y agoWhat amazes me at the end of the day is that brute-forcing seem to do much better than I initially thought it would do. Trying random stuff just sounds stupid but with enough compute and data, I guess it could overpower smart creatures like us. I agree that CodeCraft is vastly simpler than StarCraft but the idea is the same. just try random stuff(sometimes with better logic behind it) until something works and then optimize it to perfection.
- hntrader 6y agoThat randomness has to be massively constrained, though. Well over 99.9 percent of inputs are guaranteed to lead to bad results. For example if we're randomly way pointing a drone, we're almost guaranteed not to be sending it somewhere useful.
- andyljones 6y agoIn case anyone misses the links, this is twinned with two other superb posts - one about general lessons the author learned over the course of the project https://clemenswinter.com/2021/03/24/my-reinforcement-learning-learnings/ https://clemenswinter.com/2021/03/24/my-reinforcement-learni... and one history of the project https://clemenswinter.com/2021/03/24/conjuring-a-codecraft-mind/ https://clemenswinter.com/2021/03/24/conjuring-a-codecraft-m...