7 ms·
> current frontier models > Gemini 2.5 Pro, O3, Claude Sonnet 3.7 and ChatGPT 4.1 The gap in capabilities between those models which they tested, and actual c
by joefourier 19d ago
> current frontier models
> Gemini 2.5 Pro, O3, Claude Sonnet 3.7 and ChatGPT 4.1
The gap in capabilities between those models which they tested, and actual current frontier ones is enormous. I would not trust that any conclusions they made are applicable.
- deleted 19d ago[deleted]
- sigmoid10 19d agoThe actual current frontier plays somewhere around GM level. https://chessbench-ai.github.io/#leaderboard https://chessbench-ai.github.io/#leaderboard It's also worth noting that the very latest models (GPT-6 and Fable 5.1) actually play worse than their immediate predecessors, so it is likely that the labs are not benchmaxxing for this yet. If they did, I'm sure they could come up with something superior to humans. But there is probably very little demand for this compared to IT stuff.
- htrp 19d agomore like you lose intelligence in chess by maxing for coding... hence knocking back the claims of emergent intelligence
- csande17 19d agoEven if you take that website at face value, the ELO scores shown are relative to the other AI models tested, and not comparable to the ELO scores of humans who play against other humans.
- MichaelNolan 19d agoI wonder why they didn’t throw a real chess engine in there for a baseline. There are engines where you can set the elo in the settings, so it should possible to see these LLMs relative to a human 1500 rather than just relative to each other.
- shric 19d ago> so it should possible to see these LLMs relative to a human 1500 rather than just relative to each other As a 1500 elo human I can tell you that a 1500 elo chess engine doesn't play like anything like a 1500 elo human.
- traes 19d agoThis is true, but I'm not sure it matters? I was poking around at the lichess database recently and those elo calibrated bots are remarkably well calibrated, their rating variance sticks out like a sore thumb compared to human players even at similar game volumes. So it should still be a decent predictor of how good a human at that level is, even if the playstyle seems alien.
- fahrvrgnugen 18d agoI feel like every position is in the database so you could just lookup the most popular move for an arbitrary elo and that's the bot.
- shric 18d agoThat would only work for the first few (from around 10 to 20 typically depending on how close people stick to opening book) moves. Conservatively there are well over 10 to the 30 positions likely to show up in realistic games. There are of the order of 10 to the 10 or so games recorded. Thus well under one in a trillion positions are "known".
- fahrvrgnugen 15d agoIt's much smaller than that. You would be unlikely to find yourself in a novel position after 40 moves even if you were trying.
- traes 15d ago
- einszwei 19d agoProbably tells us that without labs explicitly training/tuning the models or designing the harness (with fast oracle) the LLMs aren't going to get good at those areas.
- minraws 19d agoI know HN readers and posters just read numbers and can't be bothered to read, but please read the methodology before making any claims. > About their ELO ratings from their own website: > A field-relative rating calculated within ChessBench. It compares performance among the tested models and is not a direct equivalent of a human chess rating. I am around 1600 elo in over the board I can mop up Astra Fable etc even if I give them literal infinite time and all the subagents and internet access.. Please folks at least use your AIs to read stuff before making claims. AI is not GM level, it's not even 1600, I am 1600 by using memorized openings people frequently fall for with very basic intuitions. A GM is 2600 they can beat me in under 20 moves... Why do I even scroll through this website. For a moment I truly felt fooled, but then I read like a human should. Maybe I should stop doing that will be a happier life, don't think just believe in the AGI.
- peab 19d agoWhat levels are they actually at in your experience?
- minraws 19d agoSub 1300 that's my rating in the singular official tournament I participated at. But given how easily I can crush them and how often they want to make illegal moves (btw above bench seems to use a harness that pokea the model until it gives valid moves). I would rate them around 500-800 big range but at that level it's all about if the model can recall an opening or not. If it plays good first 4-8 moves the person on the end will fumble for certain and they win. I can play good/best moves till 14-15 moves if I remember the lines and find someone who falls for it. If you could give them the lines as prompts like the best 20-30 openings then they will be around 700-800. 700 is around the rating for a human who doesn't know the tricks but can do bare minimum calculations and understands the rules thoroughly.
- Forgeties79 19d agoAs someone who used to compete for years and plays currently as a hobbyist, you’re absolutely correct. LLM’s are terrible at chess and if anyone wants to sober up their view on AI, try it yourself. Anyone who casually plays on a regular basis can beat them more often than they lose. As you said if you just know the core openings (and end games, both of which you can get a handle on with modest effort) you will generally win. Edit: reminder we had computers beating the best players in the world literally decades ago. LLM’s are remarkable tools but the current promises and expectations are ridiculous
- deleted 19d ago[deleted]
- sashank_1509 19d agoThese ratings seems very wrong, i have beaten GPT Astra max thinking in chess and my rating is close to 1500. The ratings here seem more accurate: https://chessbenchllm.onrender.com/ https://chessbenchllm.onrender.com/ GPT-6 almost never suggests an illegal move anymore while even Sol still did so time to time
- jibal 18d ago"Elo is relative to the ChessBench field." They are of course "wrong" if you don't read the faint fine print and sensibly interpret them as FIDE or similar ratings.
- sobellian 19d agoIf it's a GM then I'm Magnus Carlsen, https://lichess.org/study/27lCQqDa https://lichess.org/study/27lCQqDa.
- boesboes 18d agoDumbest thing I’ve seen today
- jibal 18d agoPlease do not post misinformation. They are not playing anywhere near GM level. "Elo is relative to the ChessBench field."
- zahlman 18d ago> The actual current frontier plays somewhere around GM level.... It's also worth noting that the very latest models (GPT-6 and Fable 5.1) actually play worse than their immediate predecessors Sorry, but I am not buying that 5.6-Sol is that much better than 5.6-Luna, which can barely be coaxed to reach the midgame with legal moves and an apparent understanding of what the position is.
- deleted 19d ago[deleted]
- sobellian 19d agoI tested both myself and a weak bot against Astra xhigh, https://lichess.org/study/27lCQqDa https://lichess.org/study/27lCQqDa. It's still pretty bad at chess, though it takes longer to devolve into illegal moves.
- phist_mcgee 18d agoThat's really cool, thanks for sharing!
- hackinthebochs 18d agoSo you weren't giving it an updated board state after every move? If you want to compare apples to apples, it should give an updated board state for each move, or you should play blindfolded.
- sobellian 18d agoI can play blindfolded. I am expert OTB (though I haven't played in a while). The game was like 18 moves of theory in the Maroczy Bind.
- Topfi 18d agoBlindfolded flex by OP aside (I can barely play when seeing the board), considering reasoning traces and their nature, if we want to be fair, a person would have to get the moves, but be allowed to write them down or draw up a board in their notepad. My working memory can barely handle five chunks, a models reasoning tokens are masses of written text in comparison.
- HarHarVeryFunny 18d agoAn LLM has been trained to do everything it does blindfolded, "only" using perfect recall of everything in it's hundreds of thousands of steps of context, and hundreds of layers of KV cache. It's a computer - it has a massive advantage over a human. The fairest apples-to-apples comparison of an LLM whose training data included chess games would be a trained human such as Magnus Carlson, who can quite happily play a dozen or more simultaneous blindfold chess games.
- aprilthird2021 18d agoThey still need supervision though
- 21asdffdsa12 18d agoSo give me a falsifiable point in time, a model you would claim succeeds at the task. One does not get to hotfix-patch updater out of the pressures of reality. Today is the day.
- deleted 18d ago[deleted]
- ares623 18d agoWell I guess this excuse is finally gonna become obsolete soon with all the "pacing" nonsense.
- user43928 18d agoThere is no need to ask. If you want to test SOTA models today, there are obviously only two: GPT-6 Astra and Fable 5.1. The models listed in the paper are from early 2025 and are no longer relevant, much less on the frontier. That Claude version is no longer available today, Gemini 2.5 Pro will be shutdown next month, and the OpenAI models are only available via the API today.
- Topfi 18d agoFortunately, a fellow commenter was so kind and did it with Astra. Didn't do that well either [0]. I'm sure GPT-7 will be super mega ASI regardless (since GPT-6 Astra already claimed AGI in the minds of Jen-Hsun, et al.)... I'll say it till there is any evidence of the contrary, LLMs are not intelligent and their capabilities solely within the realms of well tailored training data. "Just" having been trained on every rule, strategy guide and likely most games of chess on the world wide web isn't even enough for an LLM to play that game reliably. Yet the same model could code a competitive chess engine, just like a model struggling to count can write advanced maths papers. Fascinating tools, but tools nonetheless. [0] https://news.ycombinator.com/item?id=49720751 https://news.ycombinator.com/item?id=49720751
- deleted 18d ago[deleted]
- yuxi258 18d ago[dead]
- dgb23 18d agoThe gap in capabilities is mostly quantitative and not qualitative.
- RealityVoid 18d agoIs it? I am on the fence on this, but it does seem like there are some qualitative improvements between the models. Not related to your post, but a fact I keep mulling over. The fact I don't trust the current crop of LLM's enough and I consider LLM's as a tech will hit a ceiling pretty hard, it doesn't mean parallel improvement curves won't spring up out of other research that will lead to much higher capabilities than currently.
- zahlman 18d ago> but it does seem like there are some qualitative improvements between the models. It could easily seem that way, I think, in a "quantity has a quality of its own" kind of way. When you can come to the same conclusion faster, that lets you iterate more; and sometimes when you iterate you find more things.
- dezsiszabi 18d ago> Is it? Yes, it is.
- lionkor 18d agoMy read is that the improvements in quality are due to excessive use of "thinking" tokens (so, higher quantity and brute force), so I agree with that.
- zahlman 18d agoJust now I tried prompting logged-out ChatGPT (which at least claims to be 5.6-Luna) with: > Let's play a game of chess. You can take White. Please draw an ASCII rendition of the board after each move, so that we can be clear about the position. (I hoped the latter requirement would help it be "not blindfolded"; last time I used a Lichess demo board to track the position in another tab, because I have no talent for blindfold chess.) For the first couple of moves it redrew the board after each move; then it started only drawing it after my moves. And on moves 5 and 6 it dropped two minor pieces for pawns in a row without any meaningful positional advantage, and after I captured the second time, it redrew a board that was simply missing one of my pieces for no reason. It actually played better when not prompted to draw a board; in the previous session, it was spontaneously giving running commentary, which I assume was based off all the "book" theory in its training data, but it still completely fell apart at early midgame.
- titzer 18d ago[flagged]
- lirolero 18d ago[dead]
- danpalmer 18d agoSure, but installing a chess program is child/teen level general ability, and playing chess well is highly trained expert level ability. Which one are we sold AI as being?
- alpinisme 18d agoI think we are being sold AI as expert only when given tools (although that is not emphasized). The (quasi?) miracle of AI right now is that you can get an agent to accomplish the task of a team of intelligent but not exceptional humans at speeds far exceeding what the human could do. Which makes it “cheap” to throw (effectively) dozens of teams at a problem for the equivalent of hundreds of man hours. That may not be the AI of sci fi fantasy but it’s still a game changing reality.
- nutrientharvest 18d ago"Transatlantic flight will never be commercially viable, we conclude based on careful study of several aircraft designs from the 1920s"
- ponector 18d agoHow about supersonic flight?
- ggreer 18d agoI don't think that's a useful comparison. Supersonic military planes have been common for decades. We don't have supersonic passenger planes because the FAA has banned supersonic flight over land since 1973, though the agency is planning on replacing it with a noise standard. Also the original supersonic passenger aircraft were government-sponsored tech demos, not financially sustainable products. With updated laws & modern technology (cameras instead of tilting noses, more efficient engines without afterburners, lighter materials), we could have viable supersonic passenger flight.
- zeroonetwothree 18d agoTechnology keeps advancing in a domain until suddenly it doesn’t. Where are my flying cars?
- krapp 18d agoThey're called helicopters.
- freejazz 18d agoAnd what since then?
- pixl97 18d agoThis has the smell of "Why don't I have a faster horse". Why no flying cars. Because objects have mass and inertia and people are incredibly stupid. Making a flying car has been done. Making a flying car not be a weapon of mass destruction is very, very hard. Also: https://www.txdot.gov/about/newsroom/statewide/air-taxi-testing-taking-flight-in-texas.html https://www.txdot.gov/about/newsroom/statewide/air-taxi-test...
- moron4hire 18d ago> The gap in capabilities between those models which they tested, and actual current frontier ones is enormous. Same story every 4 months and yet still no breakout, winning products. I've been hearing "the AI is good now" and "it 10x's my productivity" for a over a year now. If it were true, why aren't the all-in-AI using companies 10-15 years ahead of their competition yet? Why is it still all buggy, poorly designed junk?
- orangedog 18d agoI don't get why it is hard to understand there is middle ground. People are 10x their productivity, it isn't all buggy junk, but it isn't all it is hyped up to be either. It isn't that complicated. If you hold the extreme position that there isn't any value in this, that's fine, but we're only having this discussion because these models have done what humans previously failed to do.
- freejazz 18d agoI don't think the poster disagrees with you at all. The middle ground is that there are no breakout products and that the models clearly aren't so powerful as to make these companies not produce shit code.
- autoexec 18d agoRight now AI hasn't even managed to replace all the human workers taking orders at the fast food drive thru. That's a job often performed by literal children and companies are still waiting for AI to get good enough for even that. Maybe one day it will be good enough, maybe one day it will outperform humans at such a basic task, but that day is not today. If the hype were anything close to reality, we'd see it everywhere in our lives.
- wavemode 18d agoFunny you mention this - a fast food restaurant in my town now has an LLM taking drive-thru orders. Though I highly doubt it has taken anyone's job, since most of the work is still in making, packing and handing over the food. (In fact, given the area I live in, I partially feel like the advantage they saw in it was that the LLM can speak Spanish.)