6 ms·
Computers Conquer Heads-Up Limit Hold'em
- zirkonit 12y agoMisleading headline. This is a complete -solution- for a limited poker version. Really strong, but stochastic, Hold'em algorithms have existed for a long time.
- logicalmind 12y agoSpecifically, this is for heads-up (two person) limit texas hold'em.
- Pinn2 12y agoLimited still means having a choice in what to bet. Trivial EV calculations can only tell you to bet or not, so I assume it's using something more involved.
- jib 12y agoLimit means fixed size betting. Typically 1 big blinds preflop and flop, 2 big blinds on turn and river, with a max of 4 bets total on any street. The algorithm sounds like it basically identifies the path with the least negative EV for any given action, which is essentially how humans play limit too. "If my opponent plays perfectly, what path will let me lose the least or break even"
- jchendy 12y agoIndeed, "limit" by itself is shorthand for fixed limit. There are other limit games (such as pot limit and spread limit) where you do have a choice in your bet size.
- jcr 12y ago>"The algorithm, named CFR+ by its creators, uses an improved version of a technique called counterfactual regret minimization (CFR). Past CFR algorithms have tried to solve poker by using several steps at each decision point: coming up with counterfactual values representing different game outcomes; applying a regret minimization approach to figure out the strategy leading to the best outcome; and averaging the latest strategy with all past strategies." Seems like a useful algorithm for dating sites. (only half-joking)
- javajosh 12y agoYou're not the only one to think "relationships" when the phrase "regret minimization" came up!
- MBCook 12y ago> [...] because of the huge amount of memory required; roughly 262 terabytes of memory. That’s about 268,288 times as much memory as the 1-gigabyte memory available to an iPhone 6. Wow. I can't believe that sentence made it into an article, let alone one on IEEE Spectrum. First, normal people don't know how much RAM is in a phone. Second, the numbers are reasonably identical (only differing by 2.4%) so what's the point of specifying it out that way? Why not figure out what an average amount of RAM in a new computer is (4GB?) and say "That's the RAM of about 70k new Desktops".
- StavrosK 12y ago"Unintuitive quantity 1 is over 268,288 times larger than unintuitive quantity 2!"
- coderdude 12y agoI think the people you're talking about who don't know how much RAM their phone has are the same people who don't know much RAM their desktop has. At least if you're using phones as a benchmark there is the chance a reader will be awed.
- PhantomGremlin 12y agoI was happy to see the comparison to an iPhone. At least it's from the 21st century. For many years, everything was compared to "The Library of Congress". I got real tired of reading cliches like how some fiber link "can transfer the entire 2 million books of the library of congress in 10 seconds". Or how some new storage device can "store the entire library of congress". Comparisons like that never made any sense to me. Did they mean just the text characters? With or without compression? Surely images weren't included? Doesn't the library also have stuff like maps? Why weren't they included? Etc. :-)
- vinayp10 12y agoDude, I need one of these for pokerstars
- makerops 12y agoThey exist, fwiw.
- jchendy 12y agoThis probably goes without saying, but they're definitely not allowed.
- thret 12y agoYep, you write a marginally profitable bot for Stars and then spend all your time trying to get around bot detection. I gave up although the exercise was informative. Interestingly I get pegged for a bot rather frequently anyway (well, perhaps once a month). It is EXTREMELY annoying when you have to answer a captcha while playing 40 tables, you time out everywhere no matter how fast you are.
- mmanfrin 12y ago'Conquers' in the sense it always plays +EV hands. This is exactly what 'grinders' do -- you just play the hands that have a long-term positive expected value. You play that way long enough, your risk goes to zero and you normalize out at a regular return.
- jib 12y agoA good h2h limit strategy probably plays 95% of buttons and 70% of big blinds or so, so at least 30% of the hands you play are neutral EV at best (symmetrical game, so at most 50% are +EV preflop and you play at least 80% of all hands). Only playing hands you are +EV against a full range would be exploitable as it is way too tight in heads up.
- GoatOfAplomb 12y agoI don't agree with your 50% figure. Symmetrical, sure, but remember the blind bets. Once that money is committed, each player can have a higher EV in playing the hand than in folding. (Related thought experiment - consider a game where each player has $101 in front of them before the hand starts, and then post $50 and $100 blinds. 100% of hands should be played from each position.)
- jib 12y agoYeah that's fair. It was sloppy language from me. The point I wanted to make was that you play hands where your odds of winning is <50% a large amount of the time.
- baddox 12y agoSo what happens if everyone plays that strategy? Presumably expected earnings approach zero?
- lukevdp 12y agoIf everyone plays perfectly, in the long run everyone will have the same amount of money that they started with. In most games there is a rake (fee) taken by the house, so the house would slowly take everyone's money.
- ssharp 12y agoAbysmal headline. The bot conquered heads-up, limit hold 'em, a game with next to no popularity. Conquering heads-up may be a good first step, though I'd reason that making a bot that can beat a full limit table is exponentially harder than making a bot that can beat heads-up limit. Full table (9 or 10 person) limit hold 'em has been mostly dead in casinos for well over a decade and it only had a shelf life of a few years online, even during the boom. There are better edges (for the sharks) and more excitement (for the fish) in no-limit and pot-limit games. I don't see a bot being close to be able to conquer heads-up no-limit, let alone a full table of no-limit. I suspect most heads-up limit happens in private games or the end of a limit poker tournament. It's played in scenarios that will not really be that exploitable should a person grok this bot's abilities. The article also mentions that it's opponent was another bot that was also playing a very strong strategy. That suggests it's not really able to adapt to individual play, which is essential to being at profitable at all but the lowest levels of poker.
- stevewepay 12y agoI play poker in the card rooms around the Bay Area very regularly, and limit poker is absolutely popular at the card rooms that I go to, as well as in Vegas. To say that it is dead is completely untrue.
- mmanfrin 12y agoIt's my understanding that limit is popular at places like the Oaks because No Limit was not allowed for a while, until there were means of getting around it (Oaks has 'cap no limit' for example -- can't bet above the max table buyin; Bay101 has something similar). I see the current popularity as a function more of inertia than of actual interest. A majority of the tournaments and SNGs that the Oaks runs are no limit.
- stevewepay 12y agoThis is also not true. A lot of people enjoy playing limit poker, like 3-6, because they can chase big hands much cheaper. It's a completely different strategy than no-limit, where you could potentially risk your entire stack in any hand you're in. The other thing is that poker has an unexpected (at least to me) social aspect to it. The people who play against each other are largely friendly to each other, and they play against each other every day so playing 3-6 limit lets you spend more time at the tables than 1-2 or 3-5 No Limit, etc.
- matilde 12y agoBest article so far about this: http://www.parttimepoker.com/heads-up-limit-holdem-has-been-solved http://www.parttimepoker.com/heads-up-limit-holdem-has-been-...
- Dn_Ab 12y agoThis uses a particular form of a fundamentally simple yet surprisingly powerful class of learning algorithms called regret minimization. CFR is interesting in an of itself as it specializes regret minization to play extensive form games. There are also CFR algorithms to play multiplayer and no-limit games and though the guarantees of optimality are no longer there, the players are still strong (but for now, far away from experts). The article states that this algorithm is weak to bad players but that's more an artifact of resources and training method; one advantage of minimizing regret on games instead of using linear programming is that online learning versions can adapt to exploit poor play with payoff larger than the game's value. I've also posted here before that RM solves 2 player Zero sum game more efficiently than linear programming and how it's related to boosting, portfolio optimization and as an abstraction of natural selection. http://www.pnas.org/content/111/29/10620.full http://www.pnas.org/content/111/29/10620.full
- bluecalm 12y agoIt would be nice to have more details on that CFR+ approach. If you just do what the article claims (substitute averages with regret from recent iteration) you will end up with oscillating solutions which never reach the equilibrium (I just tried it with naive CFR implementation). There is surely more to that and it would be nice to see what. As to the claim of "conquering". While there is no reason to not believe them let's see how they fare vs other near optimal AIs. There is a lot of scope for numerical mistakes when you are dealing with solutions that big and it may well be that they missed something along the way. Some other teams claimed they solved HU holdem some time ago (and without supercomputers). They compete in Alberta yearly championship every year so it will be easy to see how it goes for Cepheus.
- curun1r 12y ago> While there is no reason to not believe them let's see how they fare vs other near optimal AIs They mentioned that: Burch warned that human poker players should take the Cepheus strategy with a grain of salt. After all, Cepheus honed its strategy by playing the equivalent of a near-perfect opponent that made practically no mistakes. Certain strategies that wouldn’t work against such a powerful opponent could still prove very profitable for human poker players when exploiting the mistakes of other human players.
- jcr 12y agoThere are more details on CFR+ in the Science Magazine article [1] and the in paper [2]. These were posted by hn user 'benktbyte' in another story. [1] http://www.sciencemag.org/content/347/6218/145 http://www.sciencemag.org/content/347/6218/145 [2] https://pdf.yt/d/qv-O9AwQuV1Kjb04 https://pdf.yt/d/qv-O9AwQuV1Kjb04
- bluecalm 12y agoThank you very much. I see that they've changed how regret matching works and it that context it makes to only remember values from last iteration (or average from n last iterations). Very interesting idea. I am still unable to make it work faster than naive CFR but I am probably missing some details.
- wglb 12y agoIt would appear that this is some of the team at Alberta that "solved" checkers: http://en.wikipedia.org/wiki/Chinook_(draughts_player) http://en.wikipedia.org/wiki/Chinook_(draughts_player) The book http://www.amazon.com/One-Jump-Ahead-Jonathan-Schaeffer-ebook/dp/B002C1AN3Y/ref=sr_1_3?ie=UTF8&qid=1420764059&sr=8-3&keywords=one+jump+ahead http://www.amazon.com/One-Jump-Ahead-Jonathan-Schaeffer-eboo... One Jump Ahead details the checkers effort. I highly recommend this book. The "solved" aspect of resembles enumerating all possible outcomes.
- ikeboy 12y ago>After all, Cepheus honed its strategy by playing the equivalent of a near-perfect opponent that made practically no mistakes. So why don't you f king use that instead of making your own program? Also, how do they know that play is optimal if poker hadn't been solved before? They're comparing it to perfect play, but if we can compute that, then the problem is done anyway.
- olegbl 12y agoIt's possible they simply created another AI and gave it full information (including the opponent's hand) - essentially making a "perfect" player (who cheats). Not a useful program, but a good play-mate.
- ikeboy 12y agoOptimal play against a player who knows your hand is not optimal against one who doesn't; for example you should never bluff in this case. Your suggestion makes no sense, sorry.
- clintboxe 12y agoWhat language is it written in ?