6 ms·
I know HN readers and posters just read numbers and can't be bothered to read, but please read the methodology before making any claims. > About their ELO rati
by minraws 19d ago
I know HN readers and posters just read numbers and can't be bothered to read, but please read the methodology before making any claims.
> About their ELO ratings from their own website:
> A field-relative rating calculated within ChessBench. It compares performance among the tested models and is not a direct equivalent of a human chess rating.
I am around 1600 elo in over the board I can mop up Astra Fable etc even if I give them literal infinite time and all the subagents and internet access..
Please folks at least use your AIs to read stuff before making claims.
AI is not GM level, it's not even 1600, I am 1600 by using memorized openings people frequently fall for with very basic intuitions.
A GM is 2600 they can beat me in under 20 moves...
Why do I even scroll through this website. For a moment I truly felt fooled, but then I read like a human should.
Maybe I should stop doing that will be a happier life, don't think just believe in the AGI.
- peab 19d agoWhat levels are they actually at in your experience?
- minraws 19d agoSub 1300 that's my rating in the singular official tournament I participated at. But given how easily I can crush them and how often they want to make illegal moves (btw above bench seems to use a harness that pokea the model until it gives valid moves). I would rate them around 500-800 big range but at that level it's all about if the model can recall an opening or not. If it plays good first 4-8 moves the person on the end will fumble for certain and they win. I can play good/best moves till 14-15 moves if I remember the lines and find someone who falls for it. If you could give them the lines as prompts like the best 20-30 openings then they will be around 700-800. 700 is around the rating for a human who doesn't know the tricks but can do bare minimum calculations and understands the rules thoroughly.
- Forgeties79 19d agoAs someone who used to compete for years and plays currently as a hobbyist, you’re absolutely correct. LLM’s are terrible at chess and if anyone wants to sober up their view on AI, try it yourself. Anyone who casually plays on a regular basis can beat them more often than they lose. As you said if you just know the core openings (and end games, both of which you can get a handle on with modest effort) you will generally win. Edit: reminder we had computers beating the best players in the world literally decades ago. LLM’s are remarkable tools but the current promises and expectations are ridiculous
- little_endorian 19d agoYou can take LLMs out of opening knowledge by playing chess960, and their performance degrades significantly. I just tried playing Claude Sonnet 5 (high), and it made its first illegal move on move 5. They played 4...c6, followed by 5...Nc6, somehow forgetting about the pawn the just put on c6. (My move in between was 5. Nc3, and apparently they were trying to mirror me.)
- zug_zug 19d agoSo you can see an actual game on that website, and the play seems pretty decent to me for a while (~1700 lichess = 1300 elo) until move 28 when black throws away their queen for absolutely no reason in an incomprehensible blunder. In some ways this is reflective of the AI experience at large, sometimes shockingly competent but then also sometimes ludicrously incompetent.
- firmretention 19d agoI've always liked the analogy that talking to an LLM is like talking to a really, really smart person with a head injury.
- echelon 19d agoThe AI can write a chess bot program that will beat you. You're thinking about this the wrong way. The system is built and delivered as it is because that's how the providers make the most money. If they cared to have it perform well in chess games, you'd see a different shape and behavior. We shouldn't ask the multibillion dollar automated software generation system to play games with us any more than we should ask a Boeing's flight guidance system to do so.
- minraws 19d agoSo AGI needs to be trained on something to work well on it. Lovely reasoning we have right here. Delusion runs deep in HN circles. I say that as someone heavily invested in AI startups and projects and as someone working in the field. I think most people on HN should touch grass and find real human contact. Lmao Incredible reasoning all around here.
- echelon 19d agoI'm stating that certain folks are trying to use the software-generating product as an AGI/ASI and then complaining when it doesn't play chess very well. People are holding it wrong, deliberately or not. Some are inventing bad faith measures so they can claim AI sucks.
- minraws 19d agoThen why respond at all for the sake of responding? We all know AI can code, but the question it all stemmed from what if it's AGI or GM level in chess on it's own. You can't just back pedal from the statement that apparently being able to code a chess engine is the same as being good at chess. I can write a chess engine that beats Magnus Carlson without AI that alone neither makes me GM level or AGI or any of the other claims the above comments seem to be making?
- sdf32dsf 19d agoHe keeps posting with a particular type of tone. He definitely needs to touch grass.
- uncivilized 19d agoHN is no different than Reddit, or any social media for that matter, in that commenters pretend to read articles.
- xdavidliu 19d agothat is if it even a human commenter at all
- linkjuice4all 19d agoState-sponsored psyop meta comments aside, the models obviously continue to get better, but there is still a lot of 'guard railing' required to keep even the latest models completely on-task. The chess example is interesting because it's clearly a well-studied and established domain so the rules, strategies, and whatever else is in the training data should make yield excellent results; but clearly there is some behavior in these systems that's difficult to engineer out.
- YeGoblynQueenne 19d agoFor me the useful intuition is that LLMs haven't somehow magickally learned to implement any of the algorithms we know that we have used to make strong chess engines: alpha-beta minimax and Monte-Carlo Tree Search on the one hand, and obviously the ability to learn accurate evaluation functions by self-play. I mean we've done all this before in a task-specific fashion. It's useful to know that LLMs haven't managed to do that in the process of learning to represent the entire text on the web. On the other hand they have gotten say very good at machine translation without being trained exclusively (and I select the preceding word carefully) on machine translation. Edit: I'm saying this because there is this idea expressed by e.g. Ilya Sutskever, that in order to predict the next token accurately an LLM has to learn something about all of underlying reality. See for example this interview with Dwarkesh: https://x.com/biobootloader/status/1640512444958396416 https://x.com/biobootloader/status/1640512444958396416 Where Sutskever claims that "Predicting the next token well means you understand the underlying reality that led to the creation of that token". If that were true, we should have seen LLMs play good chess by now. There is a huge amount of data on playing chess floating around on the web in the form of algebraic chess notation and if LLMs were capable of learning the "underlying reality" of chess, they would already have. They haven't. Because they can't. What Sutskever is saying flies in the face of literally hundreds of years of statistical modelling, which is to say, building predictive models that, very explicitly, do not have to understand any "underlying reality" and only have to be good at modelling a dataset.
- Onavo 19d ago> even if I give them literal infinite time and all the subagents and internet access.. Don't use the word infinite in any CS claims. They can recreate or approximate monte Carlo tree search and it technically is still a correct solution in your framing of the problem so long they defeat you.
- automatic6131 19d agoHackerNews is Gell-Mann amnesia that refreshes on every comment on every thread.
- dmurray 19d ago> I am around 1600 elo in over the board I can mop up Astra Fable etc even if I give them literal infinite time and all the subagents and internet access. I don't believe this. You refer to "subagents", so this is not just an LLM but an LLM with some kind of agentic harness. Any reasonable harness and prompt, given internet access and appropriately prompted to succeed on this task, is more than capable of firing up Lichess or chess.com and relaying moves back to you. The free levels will be enough to beat you. A frontier model can also likely one shot a chess engine that plays at your level, again if given an environment in which it can do that. I completely believe the LLM on its own can't play a full game of chess at your level. Though I'd bet that with enough reinforcement learning it is possible to train a pure transformer architecture to do that. We just don't do it because there are other approaches that play chess much better.
- lirolero 19d ago[dead]
- YeGoblynQueenne 19d ago>> I know HN readers and posters just read numbers and can't be bothered to read, but please read the methodology before making any claims. This is unfair to HN readers all of whom but one did not post the comment you replied to. You can't just tar everyone with the same brush. There are thousands (hundreds of thousands?) of users on this site.
- minraws 19d agoHow many posts if I link that do the same thing will you agree this is the norm here. Not everything I have the time and energy to reply to. This chess one is just ridiculous claims on top of ridiculous claims all the way and 0 push back in the comments except mine. I don't even know if there is critical thought or we believe what we read/shared/etc
- dezsiszabi 19d ago50% + 1 of all comments
- YeGoblynQueenne 19d agoNo, I don't agree it's the norm. There is though a general tendency to opine with strong views on subjects posters have no expertise on. I think that's because many are software engineers (or equivalent) and they are used to being expected to "wing it" on whatever technical subject comes up. On the other hand you can always find informed comments by users who have specialist knowledge. And there's plenty of pushback on here about the chess thing besides your very valid points. EDIT: anyway if I can offer a bit of unsolicited advice, it won't do you or anyone any good to accuse everyone who doesn't agree with you of laziness, even if you can see e.g. they haven't really read an article. Just say the thing you wan to say and let them figure it out. Most people will appreciate that much better and you will feel better about yourself for acting like a mature adult. It's even in the site guidelines: Please don't comment on whether someone read an article. "Did you even read the article? It mentions that" can be shortened to "The article mentions that".
- 19d ago
- victorbjorklund 19d ago> Why do I even scroll through this website. Because other HN bring in their own experience telling us what is real and what is BS. Maybe next time it will be someone else with experience in something else that will call out BS and you will see it. I didn’t really think LLM:s are any near good in chess but I don’t play chess so don’t know what 1600 means. So you helped me by calling BS.
- thelaxiankey 19d agoI'm just dropping this all over this thread but you're unfortunately mistaken https://dynomight.net/more-chess/ https://dynomight.net/more-chess/
- freejazz 18d agoMore show and less tell would be appreciated.
- minraws 18d agoSummarizing here for my dear friends, the guy on the other end managed to fine tune a model gpt-3.5-fine-tune against stockfish vs stockfish games to perform at 1200 elo level against stockfish. I have been proved wrong I should have quit while I was ahead. /s