7 ms·
> For one, gpt-3.5-turbo-instruct rarely suggests illegal moves, even in the late game. It's claimed that this model "understands" chess, and can "reason", and
by sourcepluck 2y ago
> For one, gpt-3.5-turbo-instruct rarely suggests illegal moves, even in the late game.
It's claimed that this model "understands" chess, and can "reason", and do "actual logic" (here in the comments).
I invite anyone making that claim to find me an "advanced amateur" (as the article says of the LLM's level) chess player who ever makes an illegal move. Anyone familiar with chess can confirm that it doesn't really happen.
Is there a link to the games where the illegal moves are made?
- zarzavat 2y agoAn LLM is essentially playing blindfold chess if it just gets the moves and not the position. You have to be fairly good to never make illegal moves in blindfold.
- fmbb 2y agoDoes it not always have a list of all the moves in the game always at hand in the prompt? You have to give this human the same log of the game to refer to.
- xg15 2y agoI think even then it would still be blindfold chess, because humans do a lot of "pattern matching" on the actual board state in front of them. If you only have the moves, you have to reconstruct this board state in your head.
- pera 2y agoA chat conversation where every single move is written down and accessible at any time is not the same as blindfold chess.
- zbyforgotp 2y agoYou can make it available to the player and I suspect it wouldn’t change the outcomes.
- gwd 2y agoOK, but the LLM is still playing without a board to look at, except what's "in its head". How often would 1800 ELO chess players make illegal moves when playing only using chess notation over chat, with no board to look at? What might be interesting is to see if there was some sort of prompt the LLM could use to help itself; e.g., "After repeating the entire game up until this point, describe relevant strategic and tactical aspects of the current board state, and then choose a move." Another thing that's interesting is the 1800 ELO cut-off of the training data. If the cut-off were 2000, or 2200, would that improve the results? Or, if you included training data but labeled with the player's ELO, could you request play at a specific ELO? Being able to play against a 1400 ELO computer that made the kind of mistakes a 1400 ELO human would make would be amazing.
- wingmanjd 2y agoMaiaChess [1] supposedly plays at a specific ELO, making similar mistakes a human would make at those levels. It looks like they have 3 public bots on lichess.org: 1100, 1500, and 1900 [1] https://www.maiachess.com/ https://www.maiachess.com/
- lukeschlather 2y agoThe LLM can't refer to notes, it is just relying on its memory of what input tokens it had.
- sebzim4500 2y agoSure but I'm better than 99% of people at chess and if I was playing under those conditions there is a high chance I would make an illegal move.
- GaggiX 2y agoI can confirm that an advanced amateur can play illegal moves by playing blindfold chess as shown in this article.
- _heimdall 2y agoThis is the problem with LLM researchers all but giving up on the problem of inspecting how the LLM actually works internally. As long as the LLM is a black box, its entirely possible that (a) the LLM does reason through the rules and understands what moves are legal or (b) was trained on a large set of legal moves and therefore only learned to make legal moves. You can claim either case is the real truth, but we have absolutely no way to know because we have absolutely no way to actually understand what the LLM was "thinking".
- codeulike 2y agoHere's an article where they teach an LLM Othello and then probe its internal state to assess whether it is 'modelling' the Othello board internally https://thegradient.pub/othello/ https://thegradient.pub/othello/ Associated paper: https://arxiv.org/abs/2210.13382 https://arxiv.org/abs/2210.13382
- deleted 2y ago[deleted]
- mattmcknight 2y agoIt's weird because it is not a black box at the lowest level, we can see exactly what all of the weights are doing. It's just too complex for us to understand it. What is difficult is finding some intermediate pattern in between there which we can label with an abstraction that is compatible with human understanding. It may not exist. For example, it may be more like how our brain works to produce language than it is like a logical rule based system. We occasionally say the wrong word, skip a word, spell things wrong...violate the rules of grammar. The inputs and outputs of the model are human language, so at least there we know the system as a black box can be characterized, if not understood.
- _heimdall 2y ago> The inputs and outputs of the model are human language, so at least there we know the system as a black box can be characterized, if not understood. This is actually where the AI safety debates tend to lose. From where I sit we can't characterize the black box itself, we can only characterize the outputs themselves. More specifically, we can decide what we think the quality of the output for the given input and we can attempt to infer what might have happened in between. We really have no idea what happened in between, and though many of the "doomers" raise concerns that seem far fetched, we have absolutely no way of understanding whether they are completely off base or raising concerns of a system that just hasn't shown problems in the input/output pairs yet.
- grumpopotamus 2y agoI am an expert level chess player and I have multiple people around my level play illegal moves in classic time control games over the board. I have also watched streamers various levels above me try to play illegal moves repeatedly before realizing the UI was rejecting the move because it is illegal.
- deleted 2y ago[deleted]
- zoky 2y agoI’ve been to many USCF rated tournaments and have never once seen or even heard of anyone over the age of 8 try to play an illegal move. It may happen every now and then, but it’s exceedingly rare. LLMs, on the other hand, will gladly play the Siberian Swipe, and why not? There’s no consequence for doing so as far as they are concerned.
- Dr_Birdbrain 2y agoThere are illegal moves and there are illegal moves. There is trying to move your king five squares forward (which no amateur would ever do) and there is trying to move your King to a square controlled by an unseen piece, which can happen to somebody who is distracted or otherwise off their game. Trying to castle through check is one that occasionally happens to me (I am rated 1800 on lichess).
- dgfitz 2y agoMoving your king controlled by an unrealized opponent square is simply responded to with “check” no?
- james_marks 2y agoNo, that would break the rule that one cannot move into check
- 2y ago
- rgoulter 2y ago> I invite anyone making that claim to find me an "advanced amateur" (as the article says of the LLM's level) chess player who ever makes an illegal move. Anyone familiar with chess can confirm that it doesn't really happen. This is somewhat imprecise (or inaccurate). A quick search on YouTube for "GM illegal moves" indicates that GMs have made illegal moves often enough for there to be compilations. e.g. https://www.youtube.com/watch?v=m5WVJu154F0 https://www.youtube.com/watch?v=m5WVJu154F0 -- The Vidit vs Hikaru one is perhaps the most striking, where Vidit uses his king to attack Hikaru's king.
- zoky 2y agoIt’s exceedingly rare, though. There’s a big difference between accidentally falling to notice a move that is illegal in a complicated situation, and playing a move that may or may not be illegal just because it sounds kinda “chessy”, which is pretty much what LLMs do.
- ifdefdebug 2y agoyes but LLM illegal moves often are not chessy at all. A chessy illegal move for instance would be trying to move a rook when you don't notice that it's between your king and an attacking bishop. LLMs would often happily play Ba4 when there's no bishop anywhere near a square from where it could reach that square, or even no bishop at all. That's not chessy, that's just weird. I have to admit it's been a while since I played chatgpt so maybe it improved.
- banannaise 2y agoA bunch of these are just improper procedure: several who hit the clock before choosing a promotion piece, and one who touches a piece that cannot be moved. Even those that aren't look like rational chess moves, they just fail to notice a detail of the board state (with the possible exception of Vidit's very funny king attack, which actually might have been clock manipulation to give him more time to think with 0:01 on the clock). Whereas the LLM makes "moves" that clearly indicate no ability to play chess: moving pieces to squares well outside their legal moveset, moving pieces that aren't on the board, etc.
- mattmcknight 2y ago> I invite anyone making that claim to find me an "advanced amateur" (as the article says of the LLM's level) chess player who ever makes an illegal move. I would say the analogy is more like someone saying chess moves aloud. So, just as we all misspeak or misspell things from time to time, the model output will have an error rate.
- jeremyjh 2y agoYes, I don't even know what it means to say its 1800 strength and yet plays illegal moves frequently enough that you have to code retry logic into the test harness. Under FIDE rules after two illegal moves the game is declared lost by the player making that move. If this rule were followed, I'm wondering what its rating would be.
- famouswaffles 2y ago>Yes, I don't even know what it means to say its 1800 strength and yet plays illegal moves frequently enough that you have to code retry logic into the test harness. People are really misunderstanding things here. The one model that can actually play at lichess 1800 Elo does not need any of those and will play thousands of moves before a single illegal one. But he isn't just testing that one specific model. He is testing several models, some of which cannot reliably output legal moves (and as such, this logic is required)
- chis 2y agoI agree with others that it’s similar to blindfold chess and would also add that the AI gets no time to “think” without chain of thought like the new o1 models. So it’s equivalent to an advanced player, blindfolded, making moves off pure intuition without system 2 thought.
- bjackman 2y agoSo just because has different failure modes it doesn't count as reasoning? Is reasoning just "behaving exact like a human"? In that case the statement "LLMs can't reason" is unfalsifiable and meaningless. (Which, yeah, maybe it is). The bizarre intellectual quadrilles people dance to sustain their denial of LLM capabilities will never cease to amaze me.
- deleted 2y ago[deleted]
- hamilyon2 2y agoThe discussion in this thread is amazing. People, even renowned experts in their field make mistakes, a lot of mistakes, sometimes very costly and very obvious in retrospect. In their craft. Yet when LLM, trained on corpus of human stupidity, no less, make illegal moves in chess, our brain immediately goes: I don't make illegal moves in chess, how can computer play chess if it does? Perfect examples of metacognitive bias and general attribution error at least.
- sourcepluck 2y agoYou would be correct to be amazed if someone was arguing: "Look! It made mistakes, therefore it's definitely not reasoning!" That's certainly not what I'm saying, anyway. I was responding to the argument actually being made by many here, which is: "Look! It plays pretty poorly, but not totally crap, and it wasn't trained for playing just-above-poor chess, therefore, it understands chess and definitely is reasoning!" I find this - and much of the surrounding discussion - to be quite an amazing display of people's biases, myself. People want to believe LLMs are reasoning, and so we're treated to these merry-go-round "investigations".
- stonemetal12 2y agoIt isn't a binary does\doesn't question. It is a question of frequency and "quality" of mistakes. If it is making illegal moves 0.1% of the time then sure everybody makes mistakes. If it is 30% of the time then it isn't doing so well. If the illegal moves it tries to make are basic "pieces don't move like that" sort of errors then the predict next token isn't predicting so well. If the legality of the moves is more subtle then maybe it isn't too bad. But more than being able to make moves, if we claim it understands chess shouldn't be able to explain why it chose a move over another move?
- stefan_ 2y agoNo, my brain goes that the machine constantly suggesting "jump off now!" in between the occasional legal chess move probably isn't quite right in the head. And that the people suggesting this is all perfectly fine because we can post-hoc decide what legal moves are and are not even willing to entertain the notion that this invalidates their little experiment, well, maybe that's not the ones we want deploying this kind of thing.
- fl7305 2y ago> It's claimed that this model "understands" chess, and can "reason", and do "actual logic" (here in the comments). You can divide reasoning into three levels: 1) Can't reason - just regurgitates from memory 2) Can reason, but makes mistakes 3) Always reasons perfectly, never makes mistakes If an LLM makes mistakes, you've proven that it doesn't reason perfectly. You haven't proven that it can't reason.
- eimrine 2y agoDo you know what is to reason? LLM can't do Socratus' method, are there any other ways to reason?
- fl7305 2y agoNot sure what you mean by "LLM can't do Socratus' method"? But: Plenty of people struggle with playing along with Socrates' method. Can they not reason at all?
- eimrine 2y ago> Not sure what you mean by "LLM can't do Socratus' method"? I hope you can translate the conversation of philosophist vs chatgpt from Russian [1]. The conversation from the philosopher is built due to Socrates' method but chatgpt can not even react consistently. > Plenty of people struggle with playing along with Socrates' method. Can they not reason at all? I do not hold the opinion that chatgpt "struggles" with Socrates' method, I am clearly seing it can not use it at all even from answering side of Socrates' dualogue which is not that hard. Chatgpt can not use Socrates' method from questioning side of dialogue by design because it never asks questions. [1] https://hvylya.net/analytics/268340-dialogi-sergeya-dacyuka-s-ii-chatgpt-chast-1-kto-on-chto-mozhet-i-kak-ocenivaet-raznye-situacii https://hvylya.net/analytics/268340-dialogi-sergeya-dacyuka-...
- alain94040 2y ago> find me an "advanced amateur" (as the article says of the LLM's level) chess player who ever makes an illegal move Without a board to look at, just with the same linear text input given in the prompt? I bet a lot of amateurs would not give you legal moves. No drawing or side piece of paper allowed.
- deleted 2y ago[deleted]