3 ms·
The aim of the AI isn't to adopt to poor strategies, rather to play an approximate optimal strategy itself. It's aiming to be unexploitable, the further the oth
by hunl 10y ago
The aim of the AI isn't to adopt to poor strategies, rather to play an approximate optimal strategy itself. It's aiming to be unexploitable, the further the other players deviate from optimal, the more it wins. It's EV (expected value) comes from the other players not playing optimally, it doesn't care about exploiting individual weaknesses.
- skoutus 10y agoAI's aim is not to adopt poor strategies. It seeks to optimize, but the point is you can lead it to believe a point is global optimum, when in fact it is only local optimum. As to your point about the EV, this is why collusion can work. By colluding over a long enough horizon, the AI can believe that the average expected value to be something that it is not. If only one individual feign a weakness and the rest do not, then the strategy doesn't work.
- hunl 10y agoYou can't lead it to believe a point is an optimum, it's just responded to a bet size/check in isolation given the information it has. If you 'feign weakness' in a given spot it will just respond as optimally as possible to the bet size. For example attempting to feign weakness by betting small in a spot where your entire range should bet large is not tricking the AI, it's just passing up on EV for the players, good players are not going to play poorly in hope of tricking the bot for future mythical EV gain. Also, there is no 'colluding' in heads up poker.
- xapata 10y agoThen I'd say it's not a very good poker player.
- hunl 10y agoIf you define a 'not very good' strategy as losing at a maximum of 0, then sure. Playing optimally means the worst case scenario against any opponent would be breaking even. It doesn't have to be trained on individual playing styles, it is simply playing each spot theoretically correctly. An example, say the humans are getting to a river situation with too many bluffs for a given betsize, an exploit for the AI would be to always call. The opposite is also true, if they are bluffing too little it should always fold. The players notice that the AI has adjusted, and adjust their frequencies - now exploiting the AI. By taking an exploitative approach the AI leaves itself open to be exploited, this is not the goal. If this were rock paper scissors, the AI is doing the equivalent of always throwing each at 1/3 - even when it's opponent throws rock every time. It could switch to paper, but a thinking opponent will now switch to scissors, this will continue until we are back at equilibrium. The AI aims to play poker in this same fashion, having the correct frequencies of actions for a given range in every spot.
- xapata 10y agoA better AI should be able to fool the opponent into thinking it has thrown rock (metaphorically) so that the opponent throws paper while the AI instead throws scissors. Poker isn't about equilibrium, it's about misdirection and exploitation. When the table gets cold, you liven it up by convincing everyone to do a round of straddle.
- hunl 10y agoHeads up poker is precisely about equilibrium. Your straddle reference is also irrelevant, this is not live multiway poker. "Tricking an opponent into thinking it has metaphorically thrown rock" extrapolated into a poker example would be betting larger/smaller, calling more/less, folding more/less than is optimal in a given scenario in the hope that your opponent makes a (bigger) mistake. You're simply hoping he makes more errors than you, the AI instead choses to just make zero mistakes and let the opponents do the rest. You can see this in action for yourself in Heads up limit holdem by playing Cepheus (http://poker-play.srv.ualberta.ca http://poker-play.srv.ualberta.ca)
- xapata 10y agoYou're still thinking one hand at a time. It may be possible to confuse the opponent into permanently shifting strategy. I agree that would not happen if two equilibrium-seeking computers played each other. Since the human strategy is unknown, it is possible that equilibrium may not exist or be optimal. Even if it's two computers, if one of the computers has the possibility of choosing a non-equilibrium strategy, then again the optimal strategy may not be to seek equilibrium.