3 ms·
AI bros: the LLM beats humans at solving Navier-Stokes and some old cypher. We are close to AGI Also AI bros: LLM can’t beat an avg chess player. But that does
by diehunde 17d ago
AI bros: the LLM beats humans at solving Navier-Stokes and some old cypher. We are close to AGI
Also AI bros: LLM can’t beat an avg chess player. But that doesn’t mean anything. It doesn’t count
- hackinthebochs 17d ago>LLM can’t beat an avg chess player. Why should that matter?
- janalsncm 17d agoIf something has general intelligence it should be able to read the rules of a game and follow them. Therefore an artificial general intelligence (AGI) should be able to do this. So we have a situation where very powerful and influential people are saying we will have AGI in 6 months (if we don’t already), yet the facts on the ground are so clearly pointing in the opposite direction.
- hackinthebochs 17d agoI would bet a lot of money that Astra can follow the rules of chess (perhaps if repeated within the context window). Also, this is a different argument than what I responded to.
- minraws 17d agoI can write you a benchmark to prove it even with a heavy handed system prompt Astra will make an illegal move during the course of the games first few moves are generally ok since it's just throwing out learned moves.
- hackinthebochs 17d agoI'd genuinely like to see the results of that.
- janalsncm 17d agoI would definitely take you up on that.
- simianwords 17d agohttps://www.chessbench.org/ https://www.chessbench.org/ >GPT-6 Astra xHigh: 0.06% rejected moves
- frde_me 17d agoI wonder if I would do better as a human, maybe? Or would I happen to have one move in 1500+ that's not valid? I could see myself messing up something at some point if the board is complicated enough and trying an illegal move, perhaps if a piece somewhere would attack my king if I moved another piece. Even through I do know the rules of chess, and I have played a few games once every so often.
- Gregkion 17d agoSo we humans are not a general intelligence then? And the stuff i'm using LLMs daily is just fake? I see i see. I will see myself out of this weird discussion while I let an LLM continue doing a lot of interesting things.
- dosisking 17d ago[flagged]
- zahlman 17d ago> So we humans are not a general intelligence then? No, because we can, in fact, generally read the rules of a game and then follow them. It's actually a hobby for many of us. > And the stuff i'm using LLMs daily is just fake? This misses the point completely.
- hackinthebochs 17d ago> generally read the rules of a game and then follow them How many times do you think chess.com prevents illegal moves from being executed? Even Super GM's fall for mate-in-1's occasionally, which is functionally equivalent to missing a pin or a check. This idea that LLMs failing to only ever make legal moves undermines their intelligence doesn't pass the smell test.
- diehunde 17d agoDo you play chess ? Do you even know what an illegal move is ?
- hackinthebochs 17d agoIf you have something to contribute to the discussion, just say it
- zahlman 16d agoChess.com has to accommodate people who haven't learned the rules yet on the low end. On the high end, people are commonly playing fast enough that they're often outlining sequences of multiple "pre-moves" during the opponent's turn in order to avoid losing on time. And no, I would not agree with that functional equivalence.
- lelanthran 17d ago> Why should that matter? Because we want to use this as a replacement for humans, and the average human can learn the rules of chess without needing to see the rules explained hundreds of thousands of times in millions of games. So, yeah, it matters if a model has millions of examples of something in its training set and still cannot follow the rules.
- hackinthebochs 17d agoWe're not talking about learning the rules of chess here, but playing a competent game from just being shown the rules. Why is it so hard for people to keep track of the thread of discussion?
- ncruces 17d agoBut we are. The models can't even follow the rules: they try illegal moves all the time.
- lelanthran 17d ago> We're not talking about learning the rules of chess here, but playing a competent game from just being shown the rules. Okay, lets go with that: it's the "shown the rules" bit that we are arguing about. The argument is that a human may play maybe a dozen games after learning the rules, after which they won't be inadvertently attempting illegal moves. What we are observing with SOTA models is that, even after seeing millions of chess rules, rulebooks, actual games, etc, they still attempt illegal moves. This does not point to generalisable and adaptable intelligence, such as we see in the average human.
- hackinthebochs 17d agoThis is not good reasoning. Humans need at least dozens if not hundreds of reinforcement sessions to only make legal moves, and still occasionally fail (consider pins, discovered check, failing to respond to check). LLMs must one-shot a competent game after imbibing a mass of disconnected units of information about chess. Nothing about the two are similar. See my comment here for more: https://news.ycombinator.com/item?id=49725306 https://news.ycombinator.com/item?id=49725306
- HarHarVeryFunny 17d agoIt depends on what you are selling it as. It only matters if you are claiming it to be general purpose. If you admit that it's just a collection of narrow capabilities - whose strength is mostly confined to the 1000 or so RL environments it was post-trained in, then there is of course no expectation of it being general purpose. The AI companies seem to heavily want you to believe it is some some near human level general intelligence, so therefore pointing out all the things it can't do is very relevant.
- lostmsu 17d agoThe fact that LLMs can play chess at any level is a strong indication we are in AGI.
- recursive 17d agoCan they if they frequently make illegal moves?
- lostmsu 17d agoDo they?
- bigstrat2003 17d agoNo it isn't. Computers could play chess long before LLMs, better than LLMs can in fact. That didn't make them AGI.
- lostmsu 17d agoYou are saying "No it is not" without an argument. The fact that computer systems could play chess yet not being AGI has no relevance to LLMs' ability to play chess being AGI, because the point is about G, not I. There's little doubt about A or I parts.
- zahlman 17d agoThis is roughly comparable to observing a cat batting a ball away with its paw and taking this as a "strong indication" that cats can play any sport.
- lostmsu 17d agoYes, a good analogy. Except the cat actually follows the football rules and can beat some humans. And has no physical limitations to play other kinds of sport that you might imply.
- HarHarVeryFunny 17d agoIt would be more impressive if they could play chess (or do anything they haven't been custom RLVR trained for) by reasoning, rather than just "have a go at it" prediction which is closer to memorization. HOW you do it makes a big difference in how you should assess the capability of the thing doing it. Stockfish will trounce any LLM, and any human, at chess, so should we say that Stockfish is smarter than both?