4 ms·
I'm guessing they're training the AI independently for each player, with that in mind I wonder how good it would perform when matched against a different player
by necessity 10y ago
I'm guessing they're training the AI independently for each player, with that in mind I wonder how good it would perform when matched against a different player. Knowing your opponent's playing style is one of the main things you have to master in poker. This is specially difficult as playing style will vary as the game goes, and sometimes solely by the player's will (to trick you) and not by some information derived from the game. I can see a bot performing reasonably well on this task for a specific player despite all the difficulties, but for the general case it seems like a huge challenge (as it is for humans).
- conistonwater 10y agoThe article says they're computing (approximating) the mixed Nash equilibrium, which means it doesn't depend on the opponent's style.
- necessity 10y agoI noticed that, I just presumed they did some training.
- splonk 10y agoWithout having looked into the details of this particular bot, they're almost always trained against themselves. It would be very surprising to me if they bootstrapped with any hands from live players.
- xapata 10y agoIf it's just looking for equilibrium, it won't make any money. It also does live analysis of end-game situations, so I'm guessing it'll make use of U of Alberta's research into deviating from equilibrium to exploit leaks in the opponent's play.
- spectrum1234 10y agoIt will if the human is not playing perfectly. Which of course they aren't because they are human. The question is if the computer is playing close enough to equilibrium/optimal compared to the humans.
- xapata 10y ago> equilibrium/optimal That assumes that the Nash equilibrium is optimal. It isn't. Optimal play is to deviate from Nash equilibrium to take the most exploitative strategy against specific opponent weaknesses. This necessarily creates weakness in your own play. Optimal play sends misleading signals to the opponent. Basically, you signal "rock" so that the opponent plays "paper" but you actually play "scissors". (In this metaphor, neither "rock" nor "paper" nor "scissors" are equilibrium strategies.) For example, I might order a couple beers, straddle a few times and start egging on the rest of the players for a round of straddling. To get it going I might throw in a few blind bets. Ask folks to go all-in blind once or twice, for fun. Do that a few times and no one expects you to have a good hand. Well, except for the folks that have seen that story play out before. Except for the beers and bending the rules, the computer can do the same thing.