4 ms·
So let me see if I understand this. I don't believe it's hard to write a probabilistic program to play poker. That's enough to win against humans in 2-player.
by JaRail 7y ago
So let me see if I understand this. I don't believe it's hard to write a probabilistic program to play poker. That's enough to win against humans in 2-player.
With one AI and multiple professional human players sitting at a physical table, the humans outperform the probabilistic model because they take advantage of each other's mistakes/styles. Some players crash out faster but the winner gets ahead of the safe probabilistic style of play.
So this bot is better at the current professional player meta than the current players. In a 1v1 against a probabilistic model, it would probably also lose?
Am I understanding this properly? Or is playing the probabilistic model directly enough of a tell that it's also losing strategy? Meaning you need some variation of strategies, strategy detection, or knowledge of the meta to win?
- rightbyte 7y agoInteresting article. Too bad a don't have a subscription to read the paper. The bot played like 10 000 hands. There is no way that is enough to prove it's better or worse than the opponents. More so in no-limit where some key all-ins can turn the game up side down. The variance is higher than limit or fixed, right? I did a heads up Texas holdem fixed bot with "counter factual regret minimization" like 8 years ago from a paper I read. It had to play like 100 000 hands vs a crappy reference bot to prove it was better. Strategy detection in so short games is probably worthless. The edge is probably in seeing who are tired or drunk in paper poker.
- junar 7y agoThey mention that they use AIVAT to reduce variance. > Although poker is a game of skill, there is an extremely large luck component as well. It is common for top professionals to lose money even over the course of 10,000 hands of poker simply because of bad luck. To reduce the role of luck, we used a version of the AIVAT[1] variance reduction algorithm, which applies a baseline estimate of the value of each situation to reduce variance while still keeping the samples unbiased. For example, if the bot is dealt a really strong hand, AIVAT will subtract a baseline value from its winnings to counter the good luck. This adjustment allowed us to achieve statistically significant results with roughly 10x fewer hands than would normally be needed. [1] https://arxiv.org/abs/1612.06915 https://arxiv.org/abs/1612.06915