9 ms·
More on Dota 2
- bdz 9y agoProps for the $12k donation to OpenDota. That's really awesome! Tho I personally always preferred Dotabuff.
- literallycancer 9y agoThe features are slightly different so people usually use both.
- jsnell 9y agoDupe, more discussion at https://news.ycombinator.com/item?id=15031470 https://news.ycombinator.com/item?id=15031470
- aaron695 9y agoUnfortunately the HN title there which still isn't correct at time of writing this, distroyed a proper conversation. (Assuming the original article didn't fix their title)
- deleted 9y ago[deleted]
- omarforgotpwd 9y agoI’m by no means an expert, but I’m fascinated by the idea that a neural net playing against itself can substantially outperform a supervised learning approach with a large training data set. I mean, gathering training data and making sure it’s labeled correctly and all that is a huge hassle so if you could eliminate that step or even reduce the amount or quality of training data required that should be a big win for AI, right? Especially if doing this not only makes things easier but also improves the performance of the model.
- NamTaf 9y agoReinforcement learning isn't a new idea - I did a Berkeley-based edX course on it a few years ago now and it was not state-of-the-art to my knowledge. That had no deep aspect to it, we just generated a reinfrocement algorithm that utilised a good measure of performance (specifically, it was pacman and the score value is pretty good at that) and changed a few algorithm weighting variables at each iteration. My understanding, from talking to a ML friend this morning, is that the latest progress is taking reinforcement learning and applying deep learning approaches (nets, etc.) to it. The key becomes finding the right scoring algorithms to tweak the neural net correctly towards the desired outcome. The self-play really is the reinforcement side of things at work. How you take that 'score' and use it to correctly modify the input weightings - be them in a neural net, traditional algorithm, etc. - is the key.
- noway421 9y agoWhy wouldn't algorithm reach a local maxima when playing with itself, or even degrade over time by opening up to unknown attacks?
- mannykannot 9y ago> The key becomes finding the right scoring algorithms to tweak the neural net correctly towards the desired outcome. Does this not become something similar to supervised learning if you are scoring internal states of the game? (i.e. scoring on more than just the outcome and things that violate the rules?)
- 9y ago
- rtpg 9y agoIs there a good self-contained example of how people set up learning in ways where "the AI doesn't initially know the rules"? I've heard this many times and conceptually I get the principle, but I have a hard time understanding how you create a legitimate starting position or measurement mechanism beyond "losing/winning".
- deleted 9y ago[deleted]
- sanxiyn 9y agoThe linked article says "The bot received incentives for winning and basic metrics like health and last hits". So apart from losing/winning, losing health is bad, last hitting is good. You could add more, but apparently that's all OpenAI used.
- mannykannot 9y agoIt also says "We also separately trained the initial creep block using traditional RL techniques." I have no idea how significant that is, but it seems to be getting a fair amount of attention.
- gcp 9y agoIt's a highly specific procedure that happens before there is interaction with the opponent, so without "handing over" the understanding that having creeps on your high-ground is good, it's very hard for the learning to see through the noise and discover this.
- damnfine 9y agoExactly why this is not impressive to me. The point is to be able to learn the rules, but all I see is some of not only the rules, but the actions already prespecified in many cases. Yes its hard, and thats why humans still rule the roost.
- kensai 9y agoIt is amazing what Peter Thiel's OpenAI does! Congrats to his genius.
- lvoudour 9y agoI know it has been mentioned a lot the past few days, but since the articles keep flowing about it I'll mention it again: It's a great feat and kudos to the openai team, but it is VERY unfair for the human players who rely on a sensory interface vs a direct API connection. That's unlike chess or go where the interface isn't important. The really impressive feat will be an AI that uses the same sensory information to make decisions (and I really hope that's where the openai will head next)
- hobofan 9y agoAnd the response as I've seen it on other threads: It probably doesn't make a big difference, and will outperform humans there too, and it would be a huge waste of computing power to train it that way. I think OpenAI should show that the AI can derive (a close aproximation of) the API data from videos, but I don't think that building a closed training loop would add much value here.
- lvoudour 9y ago>And the response as I've seen it on other threads: It probably doesn't make a big difference, and will outperform humans there too, and it would be a huge waste of computing power to train it that way. Well they may be right about the "outperform" part but they are dead wrong about the waste of time/effort/energy part. I mean if (at least human-like) real-time video/audio recognition and decision making is not an impressive AI feat, I don't know what is. I'm no expert in the field but claiming that plugging into an API and crunching numbers is more important than sensory-based decision making, just doesn't sound right
- radarsat1 9y agoIt's their long-term goal: https://blog.openai.com/universe/ https://blog.openai.com/universe/
- hobofan 9y ago> sensory-based decision making You can pretty cleanly split that up into two different problems: "sensory-based data extraction" and "data-based decision making" Though I haven't worked in the field of self-driving cars, I am fairly confident that they employ a similar split: One part that takes in all the (pre-processed) data from LIDAR, cameras, etc. and maps that to a simplified model of the surroundings, and another part that makes the driving decisions based on the simplified model. Sensory->data mapping doesn't raise a lot of eyebrows anymore if you can generate as much sensory information as you want to explore all possible states, as it is possible with Dota.
- Havoc 9y agoReally cool writeup. Enjoyed that thoroughly
- kyberias 9y agoThe classic reinforcement learning -based AI (from 1992) that beats humans (maybe not top players though) in Backgammon: https://en.wikipedia.org/wiki/TD-Gammon https://en.wikipedia.org/wiki/TD-Gammon
- juskrey 9y agoAI players that use internal game calls could beat humans from the beginning of the history.
- CoffeeBob 9y agoDoes anyone else have a problem with the line, "the graph is surprisingly linear, meaning the team improved the bot exponentially over time"?
- pycal 9y agoI think what they're pointing to is that the ELO system "true skill" that dota uses is a log normal distribution. To your point I'm not sure that means that player skill improves along the distribution, but I think it does mean their probability of winning increases exponentially.
- deleted 9y ago[deleted]
- debacle 9y agoI find the mechanism for learning + the timeline far more impressive than what they accomplished. A series of 5 bots that can consistently compete with 5 humans at the 4.5k+ level would be a very impressive display of AI training.
- distances 9y agoWhat I'd like to see is them implementing AI for a strategy game that benefits from an overall vision, such as Civilization. And then selling the AI to Firaxis. Yes, Civ 6 AI still isn't anywhere close to what it should be.
- Bartweiss 9y agoAgreed. There's been a lot of criticism of their final outcome, which is understandable - it's not at all "beats pro players at DOTA". But seeing that timeline absolutely blew me away. In particular, I'm stunned that the pro players accurately assessed "Sumail will win" on the 9th, but the improvements of one day of training invalidated the assessment.
- taormina 9y agoThat's a good hero to lead off with. Soulstealer (is the HoN name, blanking on the original Dota name) is a hero with very basic mechanics. Or rather, there's a ranged, AoE ability, an ultimate that boils down to "stand in the middle and hit the button", but the rest of the mechanics boil down to "last hit lane creeps well" which is a huge Dota 2 game mechanic. And this hero does better as they succeed at last hitting.
- yurrzz 9y agoNevermore, the Shadow Fiend is the original Dota name.
- grogenaut 9y agoI agree that it's a bit of a over hype tactic in a toy situation which might be marking to maybe make people think the technique and bot are more capable than it is. But I think people are overly down on it too. It's a demonstration of a technique in a way the general public (and gamers) will understand. OpenAI isn't going to make money off of building game bots... people wouldn't watch. The human drama is a major ingredient in e-sports. But we shouldn't be down on this while we were going gaga over a lego sorter done the same way a month or so back. It looks like for some things we can almost have a plug and play ai solution. EG, like we are seeing with image classifiers, this doesn't take years of phd doctoral research and game theory to build up a world class bot. Which is what everyone used to do. This is moving some of these techniques into the "get data set, get hardware, download library, train" plug and play type solution which we're seeing more and more with in other areas like machine classification. Eg stuff anyone with a few years of experience can do, maybe not amazingly, but better than they could hand coding the solution. The problem becomes one of gathering good training sets or building an accurate simulation to train in. This means, I think, that you'll see way more of these types of ai solutions where people would have balked at a hand coded solution before. This in turn looks a lot like mobile's change to computing where things that were annoying to do on your home pc became different just because you had a camera + gps + computer + radio in your pocket. I know my company has started using classifiers a lot more for things that are kinda sliding bad user actions instead of coding up huge rules engines. We may not be as effective as a several area deep engineers writing rules and doing data analysis, but instead we have 1 engineer per problem space being about 70% as effective which is still a huge win over not solving the problems at all. The funny thing is that this bot actually pulled off the stereotypical hollywood training montage with just a few weeks of hard work it beat the best in the world. Just get some sweet rock in there and you've got it all.
- Aron 9y agoI believe biological explanations might account partially for the openAI bot outperforming human players.
- tahw 9y agoThe version they used for TI had a variety of rules that completely changed the metagame of 1v1 (no bottle, for example). Even ignoring the obvious API advantage, the match was unfair because the Pros had never trained under the constraints that the AI team brought.
- jdoliner 9y agoThey played under standard 1v1 tournament rules: http://wiki.teamliquid.net/dota2/Dota_2_Asia_Championships/2017/Solo_Tournament http://wiki.teamliquid.net/dota2/Dota_2_Asia_Championships/2...
- shopoholic 9y ago"Arteezy also played a match against our 7.5k semi-pro tester. Arteezy was winning the whole game, but our tester still managed to surprise him with a strategy he’d learned from the bot. Arteezy remarked afterwards that this was a strategy that Paparazi had used against him once and was not commonly practiced." Does anyone have a clue what this "strategy" that Paparazi used could be?
- misticdeveloper 9y agoSo it appears their bot isn't cheating with vision, and has its action speed capped to human levels. Interesting! Sad they had whitelisted item builds. I thought the whole point of a machine learnign bot was it was supposed to learn these themselves. 5v5 full game is way more complex than Starcraft. Hope OpenAI are ready.
- hfsktr 9y agoI loved the different ideas to throw the bot off like pulling creeps in a way that would not work for a human. Slacks' courier strategy was entertaining as well. I don't know enough (no more than a layperson) about AI to have any meaningful comment there. Do they need to train the bot on every hero the same way or does it only need to relearn the hero specifics (and not items/strategies)?
- hfsktr 9y agoAnother thing I noticed. Others talked about it using API etc etc so that means that you can't visually trick it by stopping an attack mid animation like you can with human players right?