3 ms·
1. It’s hard to trust a 2026 paper that’s showing results for such old models. 2. Chess seems to be a poor benchmark for generalized strategic reasoning. Peopl
by WhitneyLand 11d ago
1. It’s hard to trust a 2026 paper that’s showing results for such old models.
2. Chess seems to be a poor benchmark for generalized strategic reasoning. People who are good at it rely more on experience and deep domain expertise than on skills that generalize to make them experts at unrelated tasks.
3. The study sounds like proving humans will never fly because they don’t have wings. In reality, humans do fly, and Claude Fable would destroy any human at chess by coding a strong enough engine on the fly.
- manquer 11d ago> People who are good at it rely more on experience and deep domain expertise People are good are 1900 or 2100 above and the top ones who spend decades in the field i.e. deep expertise are well in the 2200-2700 range. A 1100 player is none of these things, they are purely relying on strategic reasoning there is a good chance they cannot name a single opening or articulate clearly why a move was appropriate. 1100 is quite low bar.
- svachalek 11d ago1100 at online speed chess or something, could be. I'm not that deep in the chess world but everyone I know that can make 1100 in official rating can name a dozen openings and most of the known tactics, and is pretty good at applying at least one opening.
- what 11d ago> Claude Fable would destroy any human at chess by coding a strong enough engine on the fly. Delusional, but then Claude fable also isn’t beating any human at chess, the engine is.
- senordevnyc 10d agoBut does that matter? If the goal for buyers of AI is “replace this knowledge worker”, how much does it matter that the model in a simple loop can’t do it, but the model with a strong general purpose harness and a little time to gather resources and knowledge to augment the harness going forward, plus tool calls, plus custom built tools, etc, can replace the knowledge worker? Probably the only thing saving many jobs from being replaced right now is that it’s hard to have a verification of correctness in the loop, so the agent can’t hill climb very easily.
- numitus 10d agoTests are often conducted under restricted conditions. For example, elementary school students aren't given calculators in math class, or during an interview, you are asked what encapsulation is and aren't allowed to use Google. The chess test effectively demonstrates the reasoning capabilities of an LLM without relying on brute force, because a human is incapable of calculating trillions of combinations yet plays chess successfully. This test is necessary because many complex problems cannot be solved by brute force, such as managing a business or playing Heroes 3. Therefore, we can make the assumption that if an LLM can play chess at a grandmaster level without brute force, it means it will be able to command an army or manage production.
- senordevnyc 10d agoPeople might care about this for chess, but no one really cares if an LLM can command an army or manage production of a business without any tools. If it can do those tasks reliably when given access to tools (including any tools it autonomously creates for itself), then that's more than sufficient. No one cares if an LLM is doing reasoning the way humans do it, as long as it can get the job done.
- paimapi 11d agoso prove it! get a public repo out there, have it play against some open source engines also I think the operative letter in AGI is the G - and if the G is short for 'variably competent savant-like hyperfocus on certain kinds of software coding and not any other general skill' then its not really G at all, is it?
- BobbyJo 11d agoI suck at chess. Are you saying I can't be intelligent?
- paimapi 11d agois that what I'm saying? or am I talking about AGI? perhaps there's some irony here to be explored when it comes to basic reading comprehension gaps
- jibal 10d agoThat's a polite way to put it. :-)
- BobbyJo 10d agoMy point was you are misunderstanding G, or at least applying it erroneously here. Being good at chess is not a generalization of any other body of knowledge, it is a rigorous set of rules. The only way to be good at chess is to practice chess, or to apply deep calculations. The latter is the model writing code. The illegal move aspect has more to do with a failure of online/in-context learning, which would support your point. I tend to think it is a byproduct of reasoning in language, which newer architectures would fix, but we shall see.
- paimapi 10d agochess is not just a rigorous set of rules, it is rules as foundation with layers of strategy on top. and so is, for example, scientific methodology or chemical interactions or virtually everything else under-the-sun that comprises human knowledge knowledge for chess is derived from memorizing strategies that have been well-defined for decades paired with in-game reasoning processes. this is not at all different from any other body of knowledge. Noble gases, laws of thermodynamics, organic chemistry just to name a few - these are all 'strategies' that define observed phenomena, analytical frameworks that trace a logical, rational set of interactions and which can predict the next for an AGI, all of this should be a cakewalk, trained as it were to surpass human capability in any and every domain [0] (thus the G for 'general' and not 'N' for 'narrow' [1]). it should be a natural at everything, infinitely adaptable on-the-fly. the whole point of AGI is that it surpasses human capabilities even at our frontiers and bleeding edge (unless you're private enterprise and you've moved the goalposts for industry [2]) currently, it's only AGI-seeming if it gets benchmaxxed enough. otherwise it sucks at what it does and then is only barely competent at tasks if paired with enough skills and tests to make it more diligent at its work. this makes sense to me - for any probabilistically trained tool, even one that you post-train and fill with nothing but the best-quality evidence, the ultimate result is the lowest-common-denominator output for your sample set. there's no natural reasoning the AI does itself to make itself better at what it does - it's all human curation and categorization of sources ingested paired with RLHF post-training that we can get the mediocre-at-chess-at-best results that we see now and the benchmaxxed scores against whatever arbitrary and pre-defined measure that's not AGI by any classical definition. that's a cool, useful, and powerful tool that makes our lives easier, much like a hammer, nail, and studs make mounting a picture frame easier than if we only had our hands alone [0] https://ischool.syracuse.edu/types-of-ai/#:~:text=General%20AI%20%28Strong%20AI https://ischool.syracuse.edu/types-of-ai/#:~:text=General%20... [1] https://aiethicslab.rutgers.edu/e-floating-buttons/weak-ai-narrow-ai/ https://aiethicslab.rutgers.edu/e-floating-buttons/weak-ai-n... [2] https://aibusiness.com/ml/what-exactly-is-artificial-general-intelligence-ask-deepmind- https://aibusiness.com/ml/what-exactly-is-artificial-general...
- carodgers 11d ago> Claude Fable would destroy any human at chess by coding a strong enough engine on the fly. A bash script can clone and build stockfish, feed in human moves, and reply. By your standard, this bash script would "destroy any human at chess." Are you interested in assessing the intelligence of the model, or the intelligence of the tools the model can use?
- nimbleal 10d agoMaybe practically it doesn’t matter? Perhaps AGI is not the model but the model plus everything it’s got access to. If we’re modelling intelligence in the way we seem to have to to have any coherent definition of AGI, it seems to me <model + everything it can access> is always going to be more “intelligent” than <model> alone.
- Planktonne 10d agoThat would mean we should consider any human with coding knowledge a chess grandmaster, which is obviously not the case.
- nimbleal 10d agoMy points is more that, while we have a strong intuition about where, as an entity, a human's boundaries are (i.e. where the person begins and ends), philosophically it' not immediately obvious that the analogy applies to the a model in the same way. Why should that be the line drawn that says this is the "thing" and this other stuff is external to the thing? It feels somewhat arbitrary. Of course this is a difficult question with humans too, hence my reliance on intuition above. We don't have the same cultural/biological framework to fall back on with AI.
- thinkharderdev 10d agoI think all this debate about whether an LLM can write (or download) a chess engine is sort of missing the point. For basically any economically valuable work there is no equivalent of a chess engine for it. If there were we wouldn't need humans or AI to begin with.
- Certhas 10d agoGood science, properly digested and presented takes time. The idea that anything other than a breathless blog post about the latest model snapshot is useless is really poisonous to proper debate on AI issues
- WhitneyLand 10d agoNot sure how that vague truism applies to this paper. Lots of papers have great results that don’t depend on the latest models. However in this case it’s problematic: - They specifically make claims about the state of “current LLMs”. o3 is not representative of this. - They ask are LLMs capable of X and arrive at a negative result. If their claim was LLM’s can write coherent sentences, and their conclusion was positive, then there would be no issue using old models because the end result would be factual. However, when you have a negative result that makes a claim about the current state of all LLMs and the ones you were using are not current, by definition it draws the whole conclusion into question.