4 ms·
The actual current frontier plays somewhere around GM level. https://chessbench-ai.github.io/#leaderboard https://chessbench-ai.github.io/#leaderboard It's al
by sigmoid10 11d ago
The actual current frontier plays somewhere around GM level.
https://chessbench-ai.github.io/#leaderboard https://chessbench-ai.github.io/#leaderboard
It's also worth noting that the very latest models (GPT-6 and Fable 5.1) actually play worse than their immediate predecessors, so it is likely that the labs are not benchmaxxing for this yet. If they did, I'm sure they could come up with something superior to humans. But there is probably very little demand for this compared to IT stuff.
- htrp 11d agomore like you lose intelligence in chess by maxing for coding... hence knocking back the claims of emergent intelligence
- csande17 11d agoEven if you take that website at face value, the ELO scores shown are relative to the other AI models tested, and not comparable to the ELO scores of humans who play against other humans.
- MichaelNolan 11d agoI wonder why they didn’t throw a real chess engine in there for a baseline. There are engines where you can set the elo in the settings, so it should possible to see these LLMs relative to a human 1500 rather than just relative to each other.
- shric 11d ago> so it should possible to see these LLMs relative to a human 1500 rather than just relative to each other As a 1500 elo human I can tell you that a 1500 elo chess engine doesn't play like anything like a 1500 elo human.
- traes 11d agoThis is true, but I'm not sure it matters? I was poking around at the lichess database recently and those elo calibrated bots are remarkably well calibrated, their rating variance sticks out like a sore thumb compared to human players even at similar game volumes. So it should still be a decent predictor of how good a human at that level is, even if the playstyle seems alien.
- fahrvrgnugen 10d agoI feel like every position is in the database so you could just lookup the most popular move for an arbitrary elo and that's the bot.
- shric 10d agoThat would only work for the first few (from around 10 to 20 typically depending on how close people stick to opening book) moves. Conservatively there are well over 10 to the 30 positions likely to show up in realistic games. There are of the order of 10 to the 10 or so games recorded. Thus well under one in a trillion positions are "known".
- fahrvrgnugen 7d agoIt's much smaller than that. You would be unlikely to find yourself in a novel position after 40 moves even if you were trying.
- traes 7d agoThis is simply blatant misinformation. If you play a game online on lichess and go to the analysis board you can find when your game becomes novel. It will be within 20 turns unless you are intentionally following a known opening. In fact it will likely become unique within 10-15 turns.
- fahrvrgnugen 6d agoIt's not my experience at all. If you find yourself in a novel position within 10-15 moves it's likely a resignable one. Edit: maybe you don't understand what i mean by novel position. I mean any position that has never been reached in the billions of lichess games, including bullet games among beginners. Also yes I will concede that it's possible to make "quiet" moves. pawn nudges that barely affect anything. If you're doing those you're not 1500 ELO. You're intentionally trying to throw wrenches and I just don't see why an ELO bot even needs to bother with nonsense like that. This is supposed to be for fun / training!
- einszwei 11d agoProbably tells us that without labs explicitly training/tuning the models or designing the harness (with fast oracle) the LLMs aren't going to get good at those areas.
- minraws 11d agoI know HN readers and posters just read numbers and can't be bothered to read, but please read the methodology before making any claims. > About their ELO ratings from their own website: > A field-relative rating calculated within ChessBench. It compares performance among the tested models and is not a direct equivalent of a human chess rating. I am around 1600 elo in over the board I can mop up Astra Fable etc even if I give them literal infinite time and all the subagents and internet access.. Please folks at least use your AIs to read stuff before making claims. AI is not GM level, it's not even 1600, I am 1600 by using memorized openings people frequently fall for with very basic intuitions. A GM is 2600 they can beat me in under 20 moves... Why do I even scroll through this website. For a moment I truly felt fooled, but then I read like a human should. Maybe I should stop doing that will be a happier life, don't think just believe in the AGI.
- peab 11d agoWhat levels are they actually at in your experience?
- minraws 11d agoSub 1300 that's my rating in the singular official tournament I participated at. But given how easily I can crush them and how often they want to make illegal moves (btw above bench seems to use a harness that pokea the model until it gives valid moves). I would rate them around 500-800 big range but at that level it's all about if the model can recall an opening or not. If it plays good first 4-8 moves the person on the end will fumble for certain and they win. I can play good/best moves till 14-15 moves if I remember the lines and find someone who falls for it. If you could give them the lines as prompts like the best 20-30 openings then they will be around 700-800. 700 is around the rating for a human who doesn't know the tricks but can do bare minimum calculations and understands the rules thoroughly.
- Forgeties79 11d agoAs someone who used to compete for years and plays currently as a hobbyist, you’re absolutely correct. LLM’s are terrible at chess and if anyone wants to sober up their view on AI, try it yourself. Anyone who casually plays on a regular basis can beat them more often than they lose. As you said if you just know the core openings (and end games, both of which you can get a handle on with modest effort) you will generally win. Edit: reminder we had computers beating the best players in the world literally decades ago. LLM’s are remarkable tools but the current promises and expectations are ridiculous
- deleted 11d ago[deleted]
- sashank_1509 11d agoThese ratings seems very wrong, i have beaten GPT Astra max thinking in chess and my rating is close to 1500. The ratings here seem more accurate: https://chessbenchllm.onrender.com/ https://chessbenchllm.onrender.com/ GPT-6 almost never suggests an illegal move anymore while even Sol still did so time to time
- jibal 10d ago"Elo is relative to the ChessBench field." They are of course "wrong" if you don't read the faint fine print and sensibly interpret them as FIDE or similar ratings.
- sobellian 11d agoIf it's a GM then I'm Magnus Carlsen, https://lichess.org/study/27lCQqDa https://lichess.org/study/27lCQqDa.
- boesboes 10d agoDumbest thing I’ve seen today
- jibal 10d agoPlease do not post misinformation. They are not playing anywhere near GM level. "Elo is relative to the ChessBench field."
- zahlman 10d ago> The actual current frontier plays somewhere around GM level.... It's also worth noting that the very latest models (GPT-6 and Fable 5.1) actually play worse than their immediate predecessors Sorry, but I am not buying that 5.6-Sol is that much better than 5.6-Luna, which can barely be coaxed to reach the midgame with legal moves and an apparent understanding of what the position is.