5 ms·
The Bitter Lesson is very popular right now. It seems true right now. It’s having its moment right now. That doesn’t actually mean it’s axiomatically true. Com
by nickysielicki 1mo ago
The Bitter Lesson is very popular right now. It seems true right now. It’s having its moment right now. That doesn’t actually mean it’s axiomatically true.
Commenter below gets it absolutely correct: stockfish, which runs on your 5 year old phone, is dramatically better at chess than Fable. Like, so much better that it’s not even remotely comparable. The theory of the Bitter Lesson, and it’s only a theory, is that LLMs could eventually outperform stockfish. It’s not true today and it remains to be seen whether it will ever be true. For now, specialized models are absolutely better at specialized tasks.
- Evidlo 1mo agoThis seems really backwards. The Bitter Lesson is all about large data-based approaches vs hand-crafted ones, it doesn't say anything about language models not trained specifically for chess. I can't find the comment you're referring to, but the latest versions of stockfish are based on neural networks trained on millions of games, so if anything the Bitter Lesson turned out true here.
- nickysielicki 1mo agoThe conclusion of the bitter lesson would be that a large language model trained on chess commentary as well as being trained on millions of chess games would outperform stockfish which is only trained on millions of chess games. There’s no evidence at this point that this is true.
- iainmerrick 1mo agoI think you have it backwards. The common mistake is to think “maybe if we use a blend of raw data and hand-crafted heuristics, we’ll get the best of both worlds!” But the bitter lesson says no, beyond a certain point it’s better just to use the data. Thinking that an LLM might be able to improve on purely “big data” machine learning seems to me to be the same incorrect idea. Its “intelligence” is no more useful than human intelligence. The LLM is based on a massive data corpus, sure, but the amount of data specifically about chess in there pales in comparison to just playing billions of games of chess.
- TwelveEyes 1mo agoAlso, training it on chess books is literally training it on human knowledge, and not the actual game, which is exactly what the bitter lesson says not to do.
- Dylan16807 1mo ago> I think you have it backwards. > maybe if we use a blend of raw data and hand-crafted heuristics I don't follow. They're suggesting giving raw chess data to the LLM, no heuristics involved.
- iainmerrick 1mo agoI was replying to this: The conclusion of the bitter lesson would be that a large language model trained on chess commentary as well as being trained on millions of chess games would outperform stockfish which is only trained on millions of chess games. If you can draw any lessons from chess commentary, I think it’s very reasonable to call it “hand-crafted heuristics.”
- Dylan16807 1mo agoHand-crafted even if you're feeding in the raw commentary? That seems like a weird way to consider it. Wouldn't that make LLMs in general "hand-crafted"? And raw games plus raw commentary is all the data you have. You can make more games but those can be fed to both stockfish and the LLM competitor. So it seems like a valid interpretation of the bitter lesson to me.
- iainmerrick 1mo agoYeah, "hand-crafted" is a bit of a stretch; I mean their value is in the human insight they contain. The key point I was trying to get at is that the human insights don't contain anything that can't be mined from vast amounts of gameplay. Every human insight can eventually be rediscovered and made rigorous by data (in chess, at least!) In the short term, those insights are useful, but in the longer term, they add nothing at all. Note also that "raw gameplay" here can mean new games -- you can generate as much data as you need, you don't need to rely on real recorded games.
- antihipocrat 1mo agoMaybe a future frontier LLM could approach the problem by first building its own stockfish, then applying the subsequent results
- catoc 1mo agoMaybe a future LLM after that could approach the problem by first simulating a human brain, then learning from the ‘human’ gameplay. Just kidding of course
- jurgenburgen 1mo agoOr maybe an LLM could just tool call stockfish and doesn’t need to have more than a basic understanding of chess. The bitter lesson seems extraordinarily wasteful on the compute side.
- nl 1mo ago> a large language model trained on chess commentary as well as being trained on millions of chess games would outperform stockfish which is only trained on millions of chess games Not really, if anything it's closer to the opposite. The Bitter Lesson essay literally has this as an example: > These researchers wanted methods based on human input to win and were disappointed when they did not.[1] and > Enormous initial efforts went into avoiding search by taking advantage of human knowledge, or of the special features of the game, but all those efforts proved irrelevant, or worse, once search was applied effectively at scale[1] The actual bitter lesson is this: > breakthrough progress eventually arrives by an opposing approach based on scaling computation by search and learning. The eventual success is tinged with bitterness, and often incompletely digested, because it is success over a favored, human-centric approach.[1] Applying to the "LLMs-for-chess" example the bitter lesson approach would be to put many, many more games into the LLM. Does this work? People have trained fairly small LLMs that are competitive Stockfish at the ELO 1500-2000 level, eg: https://github.com/kinggongzilla/chess-bot-3000 https://github.com/kinggongzilla/chess-bot-3000 This seems to be evidence that large LLMs probably don't have as much chess training data as Stockfish does. [1] http://www.incompleteideas.net/IncIdeas/BitterLesson.html http://www.incompleteideas.net/IncIdeas/BitterLesson.html
- ianjbutler 1mo ago> These researchers wanted methods based on human input to win and were disappointed when they did not.[1] This was/is basically a strawman though. Like maybe "human input winning" was desirable for chess masters but for computer science wonks? Not the point or the disappoint. It's always neats and scruffies fighting about using some kind of recognizable method (logic) instead of magic (ML). > breakthrough progress eventually arrives by an opposing approach based on scaling computation by search and learning. More to OP's point I think: nowadays when someone wants to beat you over the head with the bitter lesson, they aren't as careful to include learning and search. They want to say learning leads to intuition (magic) whereby we can avoid work (logic/search), and maybe argue or assume from there that neats and scruffies is settled. TBF, something like reasoning in latent space does resemble intuition! But the real lesson is confirmed every time we bother to check, and not very bitter for anyone. Search/learning/logic are ALL always necessary on any sufficiently difficult problems, and hybrids that interleave always outperform everything else. Stockfish being the example in this thread that different camps of absolutists would like to claim, but also all the MCTS examples, evolving examples, and new hybrids all the time. My favorite lately: https://arxiv.org/pdf/2511.08983 https://arxiv.org/pdf/2511.08983
- camuel 1mo agoIt's the exact opposite. The bitter lesson is that simply scaling training on more games—including self-play—trumps any hand-crafted human input, whether that's fine-tuning on human commentary or clever engineering tricks. Current models are just high-dimensional interpolation engines. The denser the data sampling, the more accurate the interpolation gets. Given a choice between denser sampling and anything else, denser sampling always wins. That is the bitter lesson. Computer chess is the canonical example of this.
- klipt 1mo agoBut the harness still matters. In the case of stockfish, the harness is a tree search around the neural network evaluations.
- inigyou 1mo agoDenser sampling only seems useful if the problem domain is in some way smooth - interpolatable. If you run it on a fractal problem domain you just learn more special cases. Chess is fractal.
- zarzavat 1mo agoChess is a brute force search problem. Humans are not good at chess, even a small computer can beat Magnus Carlsen. It would be better to compare models at how well they can write the code for chess engines, otherwise it's just saying that Fable is not a good CPU emulator, which is obvious.
- kmeisthax 1mo agoThe Bitter Lesson says that the only things that scale are search and learning. Stockfish is the best chess search engine we've got, and you can learn some good heuristics for chess search policy that will make time-limited chess search a lot more powerful. That's perfectly in line with the Bitter Lesson. In contrast, LLMs playing chess are relying solely on learned behavior. The inference harnesses surrounding them aren't designed to do chess things, they're designed to do autoregressive token decoding, which isn't a search process. Reasoning traces can resemble a search process, but they're far less efficient - the LLM would have to work out each legal move, test each one, calculate a score, and simulate minimax over all of that. Assuming the LLM is smart enough to even do all that. A hand-crafted approach can absolutely beat data if your approach unlocks more search and/or learning than the general solution.
- Dylan16807 1mo ago> A hand-crafted approach can absolutely beat data if your approach unlocks more search and/or learning than the general solution. Now let's look at the bitter lesson again. It says that general methods that leverage computation are ultimately the most effective, and by a large margin. That's different from just saying to leverage computation (which is how I would interpret "unlocks more search/learning"). If the lesson is "more computation wins, when sufficiently channeled" you're basically looking at a truism. Of course more computation beats less when it's used right. The bitter lesson is about abandoning specialization in order to get more computation, and while there's a couple ways where that helps with chess, there's a lot more ways where it's counterproductive. It looks like it's more true for Go than it is for chess, and that it's not universally true. It probably correlates with the state space.
- wavemode 1mo agoBoth you and the parent commenter seem to be misunderstanding the point the Bitter Lesson paper makes. The Bitter Lesson is about general-purpose algorithms vs. specialized algorithms. Historically, chess engines were programmed to look at a chess position and use positional understanding (imparted by the human programmers) to decide what the best move is. But eventually, the chess engines that actually became stronger than humans were instead programmed to just check every possible move and countermove and see which ones lead to a win. (I'm oversimplifying, but you get the point.) So even before Stockfish contained a neural network, it was considered an example of the success of the Bitter Lesson. As it applies to AI agents, the Bitter Lesson would predict that the best possible agent would simply possess A) a way to do anything it wants, B) a way to evaluate whether what it did was correct, and C) a ton of compute. Then just turn it loose on your task. (The fact that the "brain" of the agent is an LLM is kind of irrelevant - you could also imagine the brain just being a program that generates random syntactically-correct code. What the LLM achieves is that, the random generator would take millions of years whereas the LLM is much more efficient at creating plausibly-working code. This is analogous to a chess engine's pruning heuristics.) The hard part here is B. We've seen some great agentic successes when rewriting an existing project in a new language, since the agent can just use the project's prior test suite as its evaluator. But when developing a new project, you're still figuring out the finer details of how everything is supposed to work. As the old saying goes - writing a spec that perfectly describes how a program should work, is equivalent effort to just writing the program.
- PEe9bB7D 1mo agoMaybe depends on how you ask it? Directly, or let it write a chess program? I think the latter can yield way better results.
- Animats 1mo agoGood point. Dumb AIs are needed for customer service. Most of that industry is still at "press 1 for sales, 2 for billing..." and needs something that will run locally on a 1U server.
- cjkaminski 1mo agoYes, and the technology to improve the interface you described is already available to run hundreds of concurrent instances on a 1U server. The barrier to entry is getting the people who manage those systems to care enough to implement something better.
- MrDrMcCoy 1mo agoFact. My company's largest partner is CoreWeave, and convincing leadership that we could run it ourselves on partner discounted hardware for a lot less money has gone nowhere.
- Someone 1mo ago> Most of that industry is still at "press 1 for sales, 2 for billing..." and needs something that will run locally on a 1U server. Needs? Customers want something that immediately answers their question/solves their problem, but that’s far away, even ignoring the “run locally on a 1U server” and that that may not be in the company’s interest. For many companies, that support line is a cost center, not a PR mechanism. Also “Press 1 for sales, 2 for billing...” has the big advantage that it handles all accents, speech impediments, etc. Long term I think a solution where a user’s agent trained on their voice, running on their phone communicates with the support agent of a company is where we will end up, and support phone lines will cease to exist.
- mikepurvis 1mo agoBut isn't that really just about giving "front end" models more access to specialized tool libraries, which include models tuned to specific tasks? Like the first model says ah, we're being asked to code something, oh and we've been provided with some example code, let me invoke a tool call to my model the recognizes many languages, that model says that we're looking at ocaml. Okay, I better pass this off to my ocaml model which will decipher the supplied code and make a plan for what we do about the user's intent. The ocaml model recognizes that there are tests in the supplied code, let's have the special testing model have a look at the testing strategy and see how that fits in with what we just implemented, etc etc. And perhaps at the end it all gets a single pass by a god-tier model for overall sanity and congruence, but the actual work, planning, coordination, and even user interaction was done by cheaper and faster agents of much more limited capability.
- BobbyTables2 1mo agoIt’s kinda funny that your last paragraph is basically describing why sparse files, sparse matrices, etc. are used in other contexts. It really is absurd to ask programming questions to a model also trained about the lifecycle of a fruit fly. Instead of building small models from scratch, we train an enormous model and use ridiculous amounts of GPU memory. In the end, the whole thing is shoved into RAM because we don’t know where the useful parts are… We certainly would know where they were if they were just in smaller models in the first place!
- deleted 1mo ago[deleted]
- edot 1mo agoYes, but an LLM will just call stockfish if it needs to play chess … sure if you arbitrarily constrain an LLM to use no tools it’ll suck at chess. But no one is using LLMs in isolation. Even consumer-grade, bone-stock ChatGPT has tools.
- fragmede 1mo agoChatGPT does not have stockfish as a tool it can call.
- nl 1mo agoAu contraire! https://mcpmarket.com/server/stockfish https://mcpmarket.com/server/stockfish https://til.simonwillison.net/llms/mcp-in-claude-and-chatgpt#chatgpt https://til.simonwillison.net/llms/mcp-in-claude-and-chatgpt...
- edot 1mo agoYeah but it can just install it. It writes arbitrary code. It can do whatever you want it to do.
- squidbeak 1mo agoYou're making a conceptual mistake here, comparing a chess tool to its operator. Deterministic tools produce superior results compared to models in many areas, so we allow models to use tooling. The correct analogy here is Fable as a second tier player assisting a SuperGM in running stockfish, then assessing its output to identify promising variations. There might be a limit somewhere that prevents the bitter lesson being axiomatic - for instance where simulations for anything can be exhaustive - so that judgement isn't needed any more as an arbiter. But while there are problems sufficiently complex or large to require a breadth models don't currently have, greater scale and compute will continue to convert to better decision making, and the bitter lesson will remain true (true enough).