4 ms·
I know this is reductive, but at their core, LLMs really are just very fancy autocomplete. The process they go through bears no resemblance to reasoning. Grab
by actsasbuffoon 1y ago
I know this is reductive, but at their core, LLMs really are just very fancy autocomplete. The process they go through bears no resemblance to reasoning.
Grab the best model you’ve got access to (o4, Gemini Pro 2.5, Claude Sonnet 3.7, etc) and try playing chess with it. The results are astonishingly bad.
It’s not just that LLMs make poorly thought out moves. They regularly make completely illegal moves (making a knight move like a pawn, for example), hallucinate new pieces into existence, change pieces from one color to another, etc.
More words have probably been written about Chess than any other game. You can literally buy a book that’s just about chess openings. There has to be a staggering amount of text about chess in the training data set. And yet, even the best models are far worse at chess than I was in 3rd grade (and I’m not great).
They do pretty well for the first few moves, often playing well known chess openings. But once they get to the part where you need to reason, they fall apart in the most outrageous ways. I suspect this reveals something about how they deal with other logic-based tasks. I’ve wondered for a while now if they basically code through using a vast amount of training data from Stack Overflow, which is why they’re so good at helping with common error codes, and so useless when doing something novel.
These new chess experiments have been eye-opening about how incapable the best models are of basic reasoning. Unless something about that changes drastically, I think LLMs have pretty much plateaued. There will undoubtedly be some advancements, but LLMs are never going to reach AGI without reasoning, nor will they be able to do most jobs.