3 ms·
What happens when you ask those same frontier models to write a chess-playing program? I feel like, this is a huge stumbling block that many people have. They
by MattCruikshank 18d ago
What happens when you ask those same frontier models to write a chess-playing program?
I feel like, this is a huge stumbling block that many people have. They'll give a model their data, and ask it questions. I vastly prefer letting the model understand the schema, and then writing functions or programs to answer those questions. I feel like I get way, way better answers. I can have it write unit tests for those functions. I can fuzz test those functions. I can look for data that doesn't fit the schema. I can process new data way faster (and with fewer tokens). I can repeatably get the same answers from the same inputs. I can check the code into a git repo and track changes to it over time. I can share the code with other people. I can review the code. I can improve the speed of the code and get the same answers. I can review the error accumulation, and improve it. I can decide how to handle anomalies, and encode those answers.
It's really neat to see what a frontier model can do itself. No doubt.
But "play chess by hand" is a frankly awful metric. It's kind of like asking someone to take a cube root of some arbitrary decimal, in their head, with no scratch paper.
- cbolton 18d agoIt's a great test of cognitive abilities. There are many claims that current LLMs surpass humans in cognitive abilities so it's noteworthy that they underperform on that test. Letting the model execute a chess program (that it wrote) would make sense if you're measuring its economic potential, but for cognition that would be cheating just like if you let a human run a chess program. The fact that the human would have a much harder time writing a useful program is irrelevant.
- MattCruikshank 18d ago> The fact that the human would have a much harder time writing a useful program is irrelevant. Why? There's a box. You give it a problem, and it comes up with a solution. Why does it matter to you if the box is strictly a LLM, or if the LLM can write code that it executes? Even neater if the box is self-contained with a local model. You provide electricity, and it comes up with solutions. Why does it matter if it can do chess "in its head", or if it has to use scratch paper?
- lionkor 18d agoThe question is to what end? This is a benchmark task, because playing chess, or solving other well-understood problems is more of a party trick than it is useful. If you let the LLM write a chess program, which it can ONLY do because there are already so many chess programs out there, then the benchmark becomes about recall of popular program source code, not chess.
- MattCruikshank 18d agoDo you want to measure the ability of the box, or measure the ability of the box with one hand tied behind its back? More to my point, I think it's stupid to have LLMs do work that should be done by programs... programs potentially written by LLMs. I'm advising people that they should think about this distinction, themselves, when they have data and want answers.
- cbolton 18d agoNeither. As I said I want to measure cognitive abilities. Your "ability of the box" is like "economic potential" in my previous comment. If that's what you want to measure, fine. But I want a deeper understanding: what is the thing doing, how is it solving problems? I want to get a sense of its abilities that is richer than a one-dimensional scale.
- MattCruikshank 18d agoI agree that it's a fascinating to crawl inside an LLM, and also to crawl inside of a human, and try to understand the processes and limitations. Like, Phineas Gage is one of the most remarkable learning opportunities we ever had. That said, it's really weird to me when people use (and judge) LLMs one way... and won't try using them another way. Like, to judge their utility, I think we should be open to letting them write code, and use the code they produce. Otherwise, it's like judging a Chromebook without an internet connection. Like, this was one of the most dishonest ads I've ever seen: https://www.youtube.com/watch?v=gDy9AUQJ3Fg https://www.youtube.com/watch?v=gDy9AUQJ3Fg This lamp, without a working power outlet? It really doesn't do anything...
- empath75 18d ago> It's a great test of cognitive abilities. It isn't. Stockfish running on your laptop can beat every human being on earth easily at chess. It's not intelligent _at all_ in any sense that matters.
- cbolton 18d agoWell it's not a perfect test so you need a bit of care in how you use it. If you have no idea what the subject is doing, then you don't know if you're measuring cognitive ability or something else (like cheating ability, or algorithmic sophistication or whatever). But failing the test is a pretty clear sign of certain cognitive abilities being poor.
- anthonyrstevens 18d ago>> There are many claims that current LLMs surpass humans in cognitive abilities Where? By whom? This is certainly not (yet) the general consensus, as I understand it. Are you taking the most optimistic / untethered comments as the strawman against which you feel the need to argue?
- cbolton 18d agoWhy the aggressive tone and the strawman rhetoric? I never said there was a consensus. Yes I'm talking more about the "optimistic" commenters and pointing out that this chess thing is a good datum to temper their enthusiasm. What's wrong with that? Also these claims are not completely without merit, it's just that LLMs seem to excel at specific "cognitive" tasks and it's interesting to see where they fail.
- deaton 18d agoWhat would happen if you asked a human developer to write a chess-playing program?
- Kotlopou 18d agoTo go a bit off-track based on your final sentence: my high school physics teacher would do any numerical calculation that came up first in his head, as an estimate, and only then use a calculator or write on the board. Usually the estimate was within ±10-20% of the correct value even for long combinations of numbers with a bunch of decimals. Cube roots didn't come up, but square roots did. The point of that was to show the use of approximations and of having an idea how much a result should be, to guard against calculator typos and the like. I think that has some metaphorical relevance for the chess example.
- jayd16 18d agoOf course I'm a super fast runner. I can get in my car and go like 100 mph.
- well_ackshually 18d ago>What happens when you ask those same frontier models to write a chess-playing program? they shit out a carbon copy of https://github.com/official-stockfish/stockfish https://github.com/official-stockfish/stockfish that they have in their training data. Still doesn't make Fable good at playing chess.
- MattCruikshank 18d agoI feel like you're saying something as odd as "Transistors still aren't good at playing chess." I'm pretty sure Fable could write AlphaZero, which has no lineage in common with stockfish.
- well_ackshually 18d agoI'm not the one posting daily about how "AIs are going to destroy the world because of how smart they are", "humans are finished" and "we've reached super duper mega intelligence". Go see Dario and Sam about that. >I'm pretty sure Fable could write AlphaZero If course it does, the paper is open and dozens of open source implementations are in its training data already. It could write AlphaStockfish, or xx_chessmaster_2000_xx, it doesn't matter if it does: it's writing a solver: it's not good at playing chess. If tomorrow I tell you that I'm so fucking good at chess I can beat Magnus, and I show up with a laptop running stockfish, you're going to laugh me out of the room.
- MattCruikshank 18d agoSo if you need a scratchpad to solve a problem, does that mean you cannot solve that problem?
- perrygeo 18d agoAsking a model to "write a throwaway program to do X" is vastly more productive and reliable than asking "do X". Running code provides a feedback loop, the model can iteratively improve the solution instead of guess. Even if you don't read the code yourself, you have a reproducible, editable, and auditable artifact if you need it.
- numitus 18d agoI am bad in chess game by itself like 1200 ELO, but I can write Programm and win player with 2600 ELO. Does it mean I am pro chess gamer?
- MattCruikshank 18d agoDo I care if Richard Feynman was only able to do nuclear physics with the help of an abacus? Sure, a Spelling Bee is a fun thing to have. Little kids work so hard. They practice for hours. There's joy and heartbreak. Prized, sometimes. Notoriety. But in the real world, computer-assisted spelling is by far the norm. Sometimes you care about the Bee, sometimes you care about the results.
- numitus 18d agoThis is called a benchmark. We run a calculation of Pi to evaluate a computer's performance, but we don't allow the script to download a ready-made solution. When we evaluate a runner, we don't let them use a bicycle. When we evaluate a new LLM, we don't allow it to send a request to a team of programmers, so using a chess engine for an LLM is cheating
- MattCruikshank 17d agoIf an LLM writes AlphaZero, and it competes with itself, and is dominant (and beats stockfish!!!), with no book positions cribbed from its learning... The LLM has a process to beat chess. Just like, if it doesn't inherently know how to multiply 13 * 17 without using Python to do it... I don't really care. Maybe you do care. Maybe you want an LLM to be able to do work, only in its head. But I kind of can't understand the desire for that limitation... I mean, I do. But it seems ridiculously arbitrary. Like driving a car in 2nd gear and complaining that it gets terrible mileage and can't go fast enough. The Drive gear is literally right there.
- numitus 15d agoThe assumption is that if an LLM is incapable of playing chess—a game with a relatively small number of pieces, clear and simple rules, and perfect information—even after reading a hundred thousand books on chess, then it is fundamentally incapable of managing an army or a factory. This is because those scenarios involve more 'pieces,' incomplete and fuzzy information, and implicit rules that need to be deduced independently. It doesn't matter whether it has tools or not. It's simply that running tests with chess is cheap, whereas testing with an army or writing a browser from scratch is quite time-consuming and expensive.