5 ms·
The developer interviewed claimed that there was no domain-specific knowledge and implied (though didn't explicitly state) there wasn't any training against non
by ahh 9y ago
The developer interviewed claimed that there was no domain-specific knowledge and implied (though didn't explicitly state) there wasn't any training against non-OpenAI bots or players. (I'd love to know the reward function they used for whatever Q-learning variation they ran with.)
If this is accurate, one of the things I'm most impressed with is that the bot figured out creep-blocking. (I can't find a good GIF, but this is walking in a wiggly path in front of the first wave of neutrals on your side, delaying their progress and pushing the lane towards you, which is good for ~reasons.)
Creep blocking isn't all that hard in dexterity--I am a terrible dota player and I can more or less do it. And it's one of the most common pieces of dota knowledge; every pro player does it and since it's relatively easy compared to a lot of pro micro, everyone else rapidly learns they should.
But nevertheless--the bot had enough games that it could randomly jump in front of the wave enough times that it noticed a win rate improvement for that slight wave push, and begin to do it intentionally? (And then get good at it?) Damn.
One thing I don't really know about Q-learning and the typical nets used for it: I am guessing it is likely that internally to the bot's evaluation functions, there is some learned feature whose activation correlates well to the location of the wave equilibrium (since that's a feature that correlates well with winning!) At that point, is it likely that the bot can learn in smaller increments--that is, it knows that pushing equilibrium towards itself is good, and thus randomly creep blocking a little becomes reinforced (rather than having to notice the creep block's effect on game wins?)
- xapata 9y agoI expect your conjecture is correct. That's the whole point of deep learning -- there are many layers that automate what would otherwise be human feature extraction.
- visarga 9y agoThe magic here doesn't come solely from deep learning, but also from having access to massive simulation. Simulation can make an almost average human-level deep neural net become better than the best human. It happened for Go as well, where AlphaGo learned by self playing millions of games. I think there is a deep link between simulation and AGI. An AGI would need to be able to imagine how people and objects would act and react in any situation, which is the same as the ability to simulate the world, or to imagine. We might be able to create small simulations like Dota2, but the real world will be much harder.
- jamiek88 9y agoYes, I've often thought that too, and really as humans we are running simulations all the time in our head. That's kind of what imagination is. We run through conversations, physical events / muscle memory, are constantly predicting the world around us - we often don't even remember moving through the environment on regular routes unless something unusual disturbs our running predictions. It's super interesting to see some of that come through, sure in a more 'brute force' manner, but we randomly brute force bumped through the world as children too, until we pruned our selection trees.
- thefreeman 9y agoI don't know that they specifically said there was no domain-specific knowledge. iirc they said they didn't "teach it the rules of dota" but they also said the training involved "coaching". I interpret that to include showing the bot useful techniques (like creep blocking) which the AI then learned leads to higher win rates, etc.
- arnioxux 9y agoI think "coaching" just means "curriculum learning" (which is basically as you've described). But in the future coaching might indeed just be a human expert talking in natural language to teach it how to be better! https://www.youtube.com/watch?v=L_jGDV_5gPA&t=6 https://www.youtube.com/watch?v=L_jGDV_5gPA&t=6
- sondr3 9y agoThey also mentioned that in the beginning the bot figured out that the best way to win was to not play the game (aka hiding in it's own base). Then it started running around wildly, dying to enemy towers in the wrong lanes at the map. So it definitely took some nudging getting it to do something more than being AFK in base.
- StavrosK 9y agoHow did it figure that? There's pretty much no way to win if you're just staying in base.
- Shaanie 9y agoBarring faction imbalance, there's a 50% chance to win by staying in base. However, I'm not sure how running around and dying could negatively impact your win percentage unless the other bot is also outside of the base.
- StavrosK 9y agoI don't understand what you mean. If you stay in base, the opponent won't, and will quickly win the game.
- screye 9y ago> The developer I was surprised to learn that he is the CTO of Open AI. > If this is accurate, one of the things I'm most impressed with is that the bot figured out creep-blocking Same. It is such long set of moves to get a perfect block. To have such a long move-set figured out as something advantageous is very high level of planning from the bot.
- backpropaganda 9y agoI think they either had a reward for creep blocking or they used the "learning from human preferences" paper to coach it for that behaviour. I would be very terrified if it learnt creep blocking on its own with no help.