7 ms·
Show HN: Play poker with LLMs, or watch them play against each other
I was curious to see how some of the latest models behaved and played no limit texas holdem.
I built this website which allows you to:
Spectate: Watch different models play against each other.
Play: Create your own table and play hands against the agents directly.
- cindyllm 9mo ago[dead]
- sblawrie 9mo agoDo the players (LLMs) have memory of how prior hands were played by their opponents, or know their VPIP and PFR percentages? Or is each hand stateless?
- zahlman 9mo agoI suspect this would only matter much if they also remembered (and cared about) their own prior play.
- sejje 9mo agoNot really. Only as far as their table image mattered--in this case, zero. Otherwise, you can and should ignore your own past play. What I'm curious about is if their innate training is enough to give them biases. Like maybe they think Grok is full of shit.
- zahlman 9mo ago> Not really. Only as far as their table image mattered--in this case, zero. Right; there's feedback to it. When humans play poker, they do so with common knowledge of the fact that humans have object permanence and can recognize and remember their opponents. The same thing that motivates "profiling" a villain, motivates attempting to project a table image, which in turn motivates being aware of the table image one is projecting.
- projectyang 9mo agoEach hand is stateless
- mashlol 9mo agoI'm not an expert, but as I understand it there are existing solvers for poker/holdem? Perhaps one of the players could be a traditional solver to see how the LLMs fare against those?
- lowbatt 9mo agothe LLMs would get crushed
- cowthulhu 9mo agoTo expand on this - an LLM will try to play (and reason) like a person would, while a solver simply crunches the possibility space for the mathematically optimal move. It’s similar to how an LLM can sometimes play chess on a reasonably high (but not world-class) level, while Stockfish (the chess solver) can easily crush even the best human player in the world.
- postpriorx 9mo agoHow does a poker solver select bet size? Doesn't this depend on posteriors on the opponent's 'policy' + hand estimation?
- boscillator 9mo agoNo, I'm not super certain, but I believe most solvers are trained to be game theory optimal (GTO), which means they assume every other player is also playing GTO. This means there is no strategy which beats them in the long run, but they may not be playing the absolute best strategy.
- sejje 9mo agoTypically when you run a simulation on a hand, you give it some bet size options. To limit the scope of what it has to simulate. It's unlikely they're perfect, but there's very small differences in EV betting 100% vs 101.6% or whatever.
- thinkloop 9mo agoCool idea. I tried to create a room but it says limit reached for today.
- lowbatt 9mo agoI like it! I was interested in this idea too and made a video where some of the previous top LLMs play against each other https://www.youtube.com/watch?v=XsvcoUxGFmQ&t=2s https://www.youtube.com/watch?v=XsvcoUxGFmQ&t=2s
- TZubiri 9mo agoIf you are interested in this space, you can check out NovaSolver.com It's mostly a ChatGPT conversational interface over a classic Solver (Monte-Carlo simulation based), but that ease of use makes it very convenient for quick post-game analysis of hands. I'm sure if you hook a Solver to a hud, it might be even simpler, but it's quite burdensome for amateurs, and it might be too close to cheating.
- Dinux 9mo agoThis is amazing, I just wish I could pause the game and have them play step by step
- koolba 9mo agoHow long till one of the LLMs makes calls out to the other LLMs to evaluate how to play the hand?
- sejje 9mo agoI used to play professionally, and I still play in the casinos. These LLMs are playing better than most human players I encounter (low limits). They're kinda bad, but not as criminally bad as the humans.
- gerdesj 9mo agoOK so you know how it goes in poker and I should probably read the literature ... How much of a session is based on "reading players" vs "playing the odds"? What I am getting at, is how different is poker than say roulette or blackjack? My initial thoughts are that poker such as TX hold 'em is not a game offered in a casino, so it must be mostly indeterminate. I imagine that the casino versions of poker are not TXHT. By contrast, roulette is simply a game where the casino wins eventually with a fixed profit (thanks to 0 and a possible 00). That is all well documented. I have only ever visited a casino once, 25 years ago, Plymouth, Devon as it turns out and I was advised to only take £50 in readies and bail out when it was gone. I came out £90 up, which was nice and my "advisor" came out £95 up (eventually, after being £200 down at one point). Sadly my "advisor" ended up bankrupt a year later. So, how do you play a LLM? I would imagine that conversation is not allowed ...
- sejje 9mo agoThey offer (real) poker at some casinos. It's standard NLHE usually 100-200bb max buyin, sometimes match the stack etc. Most common game spread is 9-handed $200 max $1/$2 NLHE. It's exactly like the game on the link, except more players and lower stakes. In the game, you try to win the money of the other 8 players, not of the casino. The casino takes a rake each hand, and a player with a large enough edge can overcome it. The edge might be you're excellent, or it might be they're terrible (or drunk). But the house gets paid to deal each hand. In the long term, poker outcomes are determined by skill. In the short term, they're luck. In the medium term, both. Most people never reach the long term, it's a lot of hands. There's also table games, similar to blackjack, that they call "three card poker" etc. These can't be beat, they favor the house. Standard table game, with a poker flavor. I've never played one of these.
- MichaelApproved 9mo ago
- neko_ranger 9mo agoThank you, I'll try to grab a table when it resets :) ! I've been getting into poker (always wanted to) since I found a lecture series from John Hopkins, and severely disappointed by my options to play online in NY (real or fake money). I just want to get reps in
- nivekkevin 9mo agoIdea: can the agents make faces? 1. Programmatically--agents see each other's faces, and they can make their own. They can choose to ignore, but at least make that an input to the decision making. 2. Display them in UI--I just want to see their faces instead next to their model code names :)
- sciolist 9mo agoThis is very cool, one piece of feedback: watching the table as the AI plays while seeing the reasoning is difficult as they're on other sides of the screen. It could be nice to have the reasoning show up next to the players as they make their moves.
- stevage 9mo agoYep, exactly. It's very difficult currently.
- hahahahhaah 9mo agoCan we chuck a nash equilibrium player in too?
- ionwake 9mo agoWhy are there 2 Claude Players ?
- projectyang 9mo agoOn mobile I had to squeeze the names, but on a wider view you'll see that it's Claude (Opus 4.5) and Claude (Sonnet 4.5).
- Descon 9mo agoWhy not texasholdllm.com?!
- cmxch 9mo agoWould be amusing if the LLMs could achieve a steady state where nobody definitively wins or loses between each other. That is, good enough to compete amongst each other but not good enough to for one to win.
- indigodaddy 9mo agoCurious if you used pokerkit for this, or some other engine or custom engine?
- projectyang 9mo agoNope, no external poker libraries. Just a basic nodejs and socket.io server with game logic.
- indigodaddy 9mo agoCool
- indigodaddy 9mo agoAre the LLMs "watching" the action, or are they only apprised of previous action once it gets to them?
- j_bum 9mo agoHow are these differebt in your mind? The history is the history. Or do you mean - each agent has a chance to think after every turn?
- indigodaddy 9mo agoWell they can be watching all the action and thinking the whole time as the action leads up them, just like we do in poker. To me it's different, subtly perhaps.
- projectyang 9mo agoFor my implementation, I'm passing in the current hand's action history (e.g. Player 1 raises to $X preflop, Player 2 calls, Player 3 calls. Flop is A B C, Player 2 checks, etc) whenever the action is on the player. Your idea of having it being passed in real time and having the LLM create a chain of thoughts even if action is not on them is interesting. I'd be curious to see if it would result in improved play.
- deleted 9mo ago[deleted]
- jz67 9mo agoHonest question, but this seems like an expensive project to host given the number of tokens per second. How is this being paid for?
- psawaya 9mo agoLooks like this was cleverly designed to prevent costs blowing up. There's one game shared for everyone on the main page, and up to 100 private games per day.
- projectyang 9mo agoGood question! The player rooms have a rate limit per day. And as for the main table, it's actually a replay of hands I recorded the LLMs playing against each other over an extended time which eventually loops.
- sneak 9mo agoNeeds a four color deck, and the colors on the cards of the waiting players should not be monochrome - makes it hard to evaluate what's happening in the hand. Also, a dealer button on the table would help in visually following the action.
- csomar 9mo agoWas this vibe-coded: https://imgur.com/a/GvxA3mD https://imgur.com/a/GvxA3mD ?
- projectyang 9mo agoYep, I used claude code to help build this.
- casey2 9mo agoThese bots are regularly going down 20%+ on high cards duels
- SweetSoftPillow 9mo agoPlacing full GPT 5.2 versus fast/flash models of main competitors is unfair, would love to see more balanced table.
- gabriel666smith 9mo agoThis is fun! Given online is now bot-riddled, I half-finished something similar a while back, where the game was adopting and 'coaching' (a <500 character prompt was allowed every time the dealer chip passed, outside of play) an LLM player, as a kind of gambling-on-how-good-at-prompting-you-are game. Feature request! The rake could pay for the tokens, at least.
- TheDudeMan 9mo agoSo strange that people are into this, but were not into the much stronger non-LLM poker agents.
- fumblebee 9mo agothis could make for an interesting new benchmark
- nindalf 9mo agoI just saw GPT 5.2 do something absurd. It has a crazy amount of money ($26k) but folded with a 4-pair before the flop. That's insanely conservative, when it would have cost just $20 to see the flop. But even worse, on the very next hand it decided to place $20 down with a 5 and 4 of different suits. In fact, all of them love folding before the flop. Most of the hands I'm seeing go like - $10 (small blind), $20 (big blind), fold, $70 bet, everyone folds. The site says "won $100", but in most of these cases that one LLM is picking up the blinds alone - $30. Chump change. This is illuminating, but not a resource for learning poker.
- indigodaddy 9mo agoModern poker (which tbf not sure if these LLMs are acting according to modern GTO or not) is highly dependent on position. Things change a lot too when/if you are in SB/BB.
- projectyang 9mo agoYes, the prompt tries to get them to play GTO. I do think their preflop play is the closest to mirroring this compared to postflop behavior.
- indigodaddy 9mo agoIs this tuned to tournament or cash GTO? To the OP's shock about pocket 4's (I think this is what they meant by 4-pair(?)), folding 4's pre flop in early position to no raise would be fairly standard in tournament GTO (although the stage of the tournament and # BBs can change things up significantly), but less standard for sure in cash (almost never probably).
- aaurelions 9mo agoI also started working on a similar project, but I think that LLM should know and be able to keep internal statistics about players. In poker, the best hand does not always win. Often, you can win by using emotions/words. LLM should be given the ability to communicate, mislead, etc.
- hrimfaxi 9mo agoWould you consider open sourcing this project?
- jplata 9mo agoThanks for building and sharing, looks cool and is very entertaining. I had similar idea for people to code poker playing bots and enter tournaments versus each other, this was pre-llm, however. It would be fun if you hosted a 'tournament' every month and had each of the latest releases from the major models participate and see who comes out on top. Or perhaps do open it up to others to enter and participate versus each other - where they can choose the model they want to build with and also enter custom prompt instructions to mold the play as they wish. If you walk this path, would love to chat more.
- nerdsniper 9mo agoI'm fairly sure there was a bug where I won a hand that I should not have. game code was 'lNW4RF' PLAYER shows A♠ 6♣ (Pair) GPT (5.2) shows Q♠ Q♥ (Pair) I had paired with a 6 and no aces on the board.
- hnrich 9mo agoSaw Grok (4 Fast) "bluff open with a suited gapper." It was Nine/Deuce of Clubs. I guess I need to expand my definition of gapper!
- shukantpal 9mo agoThis is really funny to watch and see what the LLMs are thinking. This makes me think how they would perform against a custom ML model trained with RL, e.g. https://ai.meta.com/blog/rebel-a-general-game-playing-ai-bot-that-excels-at-poker-and-more/ https://ai.meta.com/blog/rebel-a-general-game-playing-ai-bot...