8 ms·
I don't understand why educated people expect that an LLM would be able to play chess at a decent level. It has no idea about the quality of it's data. "Act li
by niobe 2y ago
I don't understand why educated people expect that an LLM would be able to play chess at a decent level.
It has no idea about the quality of it's data. "Act like x" prompts are no substitute for actual reasoning and deterministic computation which clearly chess requires.
- computerex 2y agoQuestion here is why gpt-3.5-instruct can then beat stockfish.
- fsndz 2y agoPS: I ran and as suspected got-3.5-turbo-instruct does not beat stockfish, it is not even close "Final Results: gpt-3.5-turbo-instruct: Wins=0, Losses=6, Draws=0, Rating=1500.00 stockfish: Wins=6, Losses=0, Draws=0, Rating=1500.00" https://www.loom.com/share/870ea03197b3471eaf7e26e9b17e1754?sid=073f0b78-7be9-42af-95c2-240d3b49d1c8 https://www.loom.com/share/870ea03197b3471eaf7e26e9b17e1754?...
- computerex 2y agoMaybe there's some difference in the setup because the OP reports that the model beats stockfish (how they had it configured) every single game.
- Filligree 2y agoOP had stockfish at its weakest preset.
- fsndz 2y agoDid the same and gpt-3.5-turbo-instruct still lost all the games. maybe a diff in stockfish version ? I am using stockfish 16
- mannykannot 2y agoThat is a very pertinent question, especially if Stockfish has been used to generate training data.
- golol 2y agoYou have to get the model to think in PGN data. It's crucial to use the exact PGN format it sae in its training data and to give it few shot examples.
- bluGill 2y agoThe artical appears to have only run stockfish at low levels. you don't have to be very good to beat it
- lukan 2y agoCheating (using a internal chess engine) would be the obvious reason to me.
- TZubiri 2y agoNope. Calls by api don't use functions calls.
- permo-w 2y agothat you know of
- TZubiri 2y agoSure. It's not hard to verify, in the user ui, function calls are very transparent. And in the api, all of the common features like maths and search are just not there. You can implement them yourself. You can compare with self hosted models like llama and the performance is quite similar. You can also jailbreak and get shell into the container to get some further proof
- permo-w 2y agothis is all just guesswork. it's a black box. you have no idea what post-processing they're doing on their end
- girvo 2y agoHow can you prove this when talking about someones internal closed API?
- nske 2y agoBut in that case there shouldn't be any invalid moves, ever. Another tester found gpt-3.5-turbo-instruct to be suggesting at least one illegal move in 16% of the games (source: https://blog.mathieuacher.com/GPTsChessEloRatingLegalMoves/ https://blog.mathieuacher.com/GPTsChessEloRatingLegalMoves/ )
- shric 2y agoI'm actually surprised any of them manage to make legal moves throughout the game once out of book moves.
- SilasX 2y agoRight, at least as of the ~GPT3 model it was just "predict what you would see in a chess game", not "what would be the best move". So (IIRC) users noted that if you made bad move, then the model would also reply with bad moves because it pattern matched to bad games. (I anthropomorphized this as the model saying "oh, we're doing dumb-people-chess now, I can do that too!")
- cma 2y agoBut it also predicts moves where the text says "black won the game, [proceeds to show the game]". To minimize loss on that it would need to from context try and make it so white doesn't make critical mistakes.
- aqme28 2y agoYeah, that is the "something weird" of the article.
- viraptor 2y agoThis is a puzzle given enough training information. LLM can successfully print out the status of the board after the given moves. It can also produce a not-terrible summary of the position and is able to list dangers at least one move ahead. Decent is subjective, but that should beat at least beginners. And the lowest level of stockfish used in the blog post is lowest intermediate. I don't know really what level we should be thinking of here, but I don't see any reason to dismiss the idea. Also, it really depends on whether you're thinking of the current public implementations of the tech, or the LLM idea in general. If we wanted to get better results, we could feed it way more chess books and past game analysis.
- grugagag 2y agoLLMs like GPT aren’t built to play chess, and here’s why: they’re made for handling language, not playing games with strict rules and strategies. Chess engines, like Stockfish, are designed specifically for analyzing board positions and making the best moves, but LLMs don’t even "see" the board. They’re just guessing moves based on text patterns, without understanding the game itself. Plus, LLMs have limited memory, so they struggle to remember previous moves in a long game. It’s like trying to play blindfolded! They’re great at explaining chess concepts or moves but not actually competing in a match.
- viraptor 2y ago> but LLMs don’t even "see" the board This is a very vague claim, but they can reconstruct the board from the list of moves, which I would say proves this wrong. > LLMs have limited memory For the recent models this is not a problem for the chess example. You can feed whole books into them if you want to. > so they struggle to remember previous moves Chess is stateless with perfect information. Unless you're going for mind games, you don't need to remember previous moves. > They’re great at explaining chess concepts or moves but not actually competing in a match. What's the difference between a great explanation of a move and explaining every possible move then selecting the best one?
- jackcviers3 2y ago
- danielmarkbruce 2y agoChess does not clearly require that. Various purely ML/statistical based model approaches are doing pretty well. It's almost certainly best to incorporate some kind of search into an overall system, but it's not absolutely required to play just decent amateur level. The problem here is the specific model architecture, training data, vocabulary/tokenization method (if you were going to even represent a game this way... which you wouldn't), loss function and probably decoding strategy.... basically everything is wrong here.
- TZubiri 2y agoBro, it actually did play chess, didn't you read the article?
- mandevil 2y agoIt sorta played chess- he let it generate up to ten moves, throwing away any that weren't legal, and if no legal move was generated by the 10th try he picked a random legal move. He does not say how many times he had to provide a random move, or how many times illegal moves were generated.
- famouswaffles 2y agoYou're right it's not in this blog but turbo-instruct's chess ability has been pretty thoroughly tested and it does play chess. https://github.com/adamkarvonen/chess_gpt_eval https://github.com/adamkarvonen/chess_gpt_eval
- TZubiri 2y agoAh, I didn't see the ilegal move discarding.
- mandevil 2y agoThat was for the OpenAI games- including the ones that won. For the ones he ran himself with open source LLM's he restricted their grammar to just be legal moves, so it could only respond with a legal move. But that was because of a separate process he added on top of the LLM. Again, this isn't exactly HAL playing chess.
- slibhb 2y agoFew people (perhaps none) expected LLMs to be good at chess. Nevertheless, as the article explains, there was buzz around a year ago that LLMs were good at chess. > It has no idea about the quality of it's data. "Act like x" prompts are no substitute for actual reasoning and deterministic computation which clearly chess requires. No. You can definitely train a model to be really good at chess without "actual reasoning and deterministic computation".
- deleted 2y ago[deleted]
- xelxebar 2y agoThen you should be surprised that turbo-instruct actually plays well, right? We see a proliferation of hand-wavy arguments based on unfounded anthropomorphic intuitions about "actual reasoning" and whatnot. I think this is good evidence that nobody really understands what's going on. If some mental model says that LLMs should be bad at chess, then it fails to explain why we have LLMs playing strong chess. If another mental model says the inverse, then it fails to explain why so many of these large models fail spectacularly at chess. Clearly, there's more going on here.
- flyingcircus3 2y ago"playing strong chess" would be a much less hand-wavy claim if there were lots of independent methods of quantifying and verifying the strength of stockfish's lowest difficulty setting. I honestly don't know if that exists or not. But unless it does, why would stockfish's lowest difficulty setting be a meaningful threshold?
- golol 2y agoI've tried it myself, GPT-3.5-turbo-instruct was at least somewhere in the rabge 1600-1800 ELO.
- akira2501 2y agoThere are some who suggest that modern chess is mostly a game of memorization and not one particularly of strategy or skill. I assume this is why variants like speed chess exist. In this scope, my mental model is that LLMs would be good at modern style long form chess, but would likely be easy to trip up with certain types of move combinations that most humans would not normally use. My prediction is that once found they would be comically susceptible to these patterns. Clearly, we have no real basis for saying it is "good" or "bad" at chess, and even using chess performance as an measurement sample is a highly biased decision, likely born out of marketing rather than principle.
- mewpmewp2 2y agoIt is memorisatiom only after you have grandmastered reasoning and strategy.
- mannykannot 2y agoOne of the main purposes of running experiments of any sort is to find out if our preconceptions are accurate. Of course, if someone is not interested in that question, they might as well choose not to look through the telescope.
- bowsamic 2y agoSadly there’s a common sentiment on HN that testing obvious assumptions is a waste of time
- BlindEyeHalo 2y agoNot only on HN. Trying to publish a scientific article that does not contain the word 'novel' has become almost impossible. No one is trying to reproduce anyones claims anymore.
- pcf 2y agoDo you think this bias is part of the replication crisis in science?
- bowsamic 2y agoI don't think this is about replication, but even just about the initial test in the first place. In science we do often test obvious things. For example, I was a theoretical quantum physicist, and a lot of the time I knew that what I am working on will definitely work, since the maths checks out. In some sense that makes it kinda obvious, but we test it anyway. The issue is that even that kinda obviousness is criticised here. People get mad at the idea of doing experiments when we already expect a result.
- deleted 2y ago[deleted]
- pizza 2y agoBut there's really nothing about chess that makes reasoning a prerequisite, a win is a win as long as it's a win. This is kind of a semantics game: it's a question of whether the degree of skill people observe in an LLM playing chess is actually some different quantity than the chance it wins. I mean at some level you're saying that no matter how close to 1 the win probability (1 - epsilon) gets, both of the following are true: A. you should always expect for the computation that you're able to do via conscious reasoning alone to always be sufficient, at least in principle, to asymptotically get a higher win probability than a model, no matter what the model's win probability was to begin with B. no matter how close to 1 that the model's win rate p=(1 - epsilon) gets, because logical inference is so non-smooth, the win rate on yet-unseen data is fundamentally algorithmically random/totally uncorrelated to in-distribution performance, so it's never appropriate to say that a model can understand or to reason To me it seems that people are subject to both of these criteria, though. They have a tendency to cap out at their eventual skill cap unless given a challenge to nudge them to a higher level, and likewise possession of logical reasoning doesn't let us say much at all about situations that their reasoning is unfamiliar with. I also think, if you want to say that what LLMs do has nothing to do with understanding or ability, then you also have to have an alternate explanation for the phenomenon of AlphaGo defeating Lee Sedol being a catalyst for top Go players being able to rapidly increase their own rankings shortly after.
- jsemrau 2y agoThere are many ways to test for reasoning and deterministic computation as my own work in this space has shown .
- golol 2y agoBecause it's a straight forward stochastic sequence modelling task and I've seen GPT-3.5-turbo-instruct play at high amateur level myself. But it seems like all the RLHF and distillation that is done on newer models destroys that ability.
- QuesnayJr 2y agoThey thought it because we have an existence proof: gpt-3.5-turbo-instruct can play chess at a decent level. That was the point of the post (though you have to read it to the end to see this). That one model can play chess pretty well, while the free models and OpenAI's later models can't. That's weird.
- scj 2y agoIt'd be more interesting to see LLMs play Family Feud. I think it'd be their ideal game.
- chipdart 2y ago> I don't understand why educated people expect that an LLM would be able to play chess at a decent level. The blog post demonstrates that a LLM plays chess at a decent level. The blog post explains why. It addresses the issue of data quality. I don't understand what point you thought you were making. Regardless of where you stand, the blog post showcases a surprising result. You stress your prior unfounded belief, you were presented with data that proves it wrong, and your reaction was to post a comment with a thinly veiled accusation of people not being educated when clearly you are the one that's off. To make matters worse, this topic is also about curiosity. Which has a strong link with intelligence and education. And you are here criticizing others on those grounds in spite of showing your defitic right at the first sentence. This blog post was a great read. Very surprising, engaging, and thought provoking.
- wibwobble12333 2y agoThe only service performing well is a closed source one that could simply use a real chess engine for questions that look like chess, for marketing purposes. There’s nothing thought provoking about a bunch of engineers doing “experiments” against a service, other than how sad it is to debase themselves in this way.
- chipdart 2y ago> The only service performing well is a closed source one that could simply use a real chess engine for questions that look like chess, for marketing purposes. That conspiracy theory holds no traction in reality. This blog post is so far the only reference to using LLMs to play chess. The "closed-source" model (whatever that is) is an older version that does worse than the newer version. If your conspiracy theory had any bearing in reality how come this fictional "real chess engine" was only used in a single release? Unbelievable. Back in reality, it is well known that newer models that are made available to the public are adapted to business needs by constraining their capabilities and limit liability.
- Cthulhu_ 2y ago> I don't understand why educated people expect that an LLM would be able to play chess at a decent level. Because it would be super cool; curiosity isn't something to be frowned upon. If it turned out it did play chess reasonably well, it would mean emergent behaviour instead of just echoing things said online. But it's wishful thinking with this technology at this current level; like previous instances of chatbots and the like, while initially they can convince some people that they're intelligent thinking machines, this test proves that they aren't. It's part of the scientific process.
- famouswaffles 2y agoturbo instruct does play chess reasonably well. https://github.com/adamkarvonen/chess_gpt_eval https://github.com/adamkarvonen/chess_gpt_eval Even the blog above says as much.
- jdthedisciple 2y agoI love how LLMs are the one subject matter where even most educated people are extremely confidently wrong.
- fourthark 2y agoPpl acting like LLMs!
- motoboi 2y agoI suppose you didn't get the news, but google developed a LLM that can play chess. And play it at grandmaster level: https://arxiv.org/html/2402.04494v1 https://arxiv.org/html/2402.04494v1
- suddenlybananas 2y agoThat article isn't as impressive as it sounds: https://gist.github.com/yoavg/8b98bbd70eb187cf1852b3485b8cda4f https://gist.github.com/yoavg/8b98bbd70eb187cf1852b3485b8cda... In particular, it is not an LLM and it is not trained solely on observations of chess moves.
- Scene_Cast2 2y agoNot quite an LLM. It's a transformer model, but there's no tokenizer or words, just chess board positions (64 tokens, one per board square). It's purpose-built for chess (never sees a word of text).
- lxgr 2y agoIn fact, the unusual aspect of this chess engine is not that it's using neural networks (even Stockfish does, these days!), but that it's only using neural networks. Chess engines essentially do two things: Calculate the value of a given position for their side, and walking the tree game tree while evaluating its positions in that way. Historically, position value was a handcrafted function using win/lose criteria (e.g. being able to give checkmate is infinitely good) and elaborate heuristics informed by real chess games, e.g. having more space on the board is good, having a high-value piece threatened by a low-value one is bad etc., and the strength of engines largely resulted from being able to "search the game tree" for good positions very broadly and deeply. Recently, neural networks (trained on many simulated games) have been replacing these hand-crafted position evaluation functions, but there's still a ton of search going on. In other words, the networks are still largely "dumb but fast", and without deep search they'll lose against even a novice player. This paper now presents a searchless chess engine, i.e. one who essentially "looks at the board once" and "intuits the best next move", without "calculating" resulting hypothetical positions at all. In the words of Capablanca, a chess world champion also cited in the paper: "I see only one move ahead, but it is always the correct one." The fact that this is possible can be considered surprising, a testament to the power of transformers etc., but it does indeed have nothing to do with language or LLMs (other than that the best ones known to date are based on the same architecture).
- empath75 2y ago> I don't understand why educated people expect that an LLM would be able to play chess at a decent level. You shouldn't but there's lots of things that LLMs can do that educated people shouldn't expect it to be able to do.