3 ms·
It's a great test of cognitive abilities. There are many claims that current LLMs surpass humans in cognitive abilities so it's noteworthy that they underperfor
by cbolton 19d ago
It's a great test of cognitive abilities. There are many claims that current LLMs surpass humans in cognitive abilities so it's noteworthy that they underperform on that test.
Letting the model execute a chess program (that it wrote) would make sense if you're measuring its economic potential, but for cognition that would be cheating just like if you let a human run a chess program. The fact that the human would have a much harder time writing a useful program is irrelevant.
- MattCruikshank 19d ago> The fact that the human would have a much harder time writing a useful program is irrelevant. Why? There's a box. You give it a problem, and it comes up with a solution. Why does it matter to you if the box is strictly a LLM, or if the LLM can write code that it executes? Even neater if the box is self-contained with a local model. You provide electricity, and it comes up with solutions. Why does it matter if it can do chess "in its head", or if it has to use scratch paper?
- lionkor 19d agoThe question is to what end? This is a benchmark task, because playing chess, or solving other well-understood problems is more of a party trick than it is useful. If you let the LLM write a chess program, which it can ONLY do because there are already so many chess programs out there, then the benchmark becomes about recall of popular program source code, not chess.
- MattCruikshank 19d agoDo you want to measure the ability of the box, or measure the ability of the box with one hand tied behind its back? More to my point, I think it's stupid to have LLMs do work that should be done by programs... programs potentially written by LLMs. I'm advising people that they should think about this distinction, themselves, when they have data and want answers.
- cbolton 19d agoNeither. As I said I want to measure cognitive abilities. Your "ability of the box" is like "economic potential" in my previous comment. If that's what you want to measure, fine. But I want a deeper understanding: what is the thing doing, how is it solving problems? I want to get a sense of its abilities that is richer than a one-dimensional scale.
- MattCruikshank 19d agoI agree that it's a fascinating to crawl inside an LLM, and also to crawl inside of a human, and try to understand the processes and limitations. Like, Phineas Gage is one of the most remarkable learning opportunities we ever had. That said, it's really weird to me when people use (and judge) LLMs one way... and won't try using them another way. Like, to judge their utility, I think we should be open to letting them write code, and use the code they produce. Otherwise, it's like judging a Chromebook without an internet connection. Like, this was one of the most dishonest ads I've ever seen: https://www.youtube.com/watch?v=gDy9AUQJ3Fg https://www.youtube.com/watch?v=gDy9AUQJ3Fg This lamp, without a working power outlet? It really doesn't do anything...
- cbolton 19d agoI completely agree.
- deleted 19d ago[deleted]
- empath75 19d ago> It's a great test of cognitive abilities. It isn't. Stockfish running on your laptop can beat every human being on earth easily at chess. It's not intelligent _at all_ in any sense that matters.
- cbolton 19d agoWell it's not a perfect test so you need a bit of care in how you use it. If you have no idea what the subject is doing, then you don't know if you're measuring cognitive ability or something else (like cheating ability, or algorithmic sophistication or whatever). But failing the test is a pretty clear sign of certain cognitive abilities being poor.
- anthonyrstevens 19d ago>> There are many claims that current LLMs surpass humans in cognitive abilities Where? By whom? This is certainly not (yet) the general consensus, as I understand it. Are you taking the most optimistic / untethered comments as the strawman against which you feel the need to argue?
- cbolton 19d agoWhy the aggressive tone and the strawman rhetoric? I never said there was a consensus. Yes I'm talking more about the "optimistic" commenters and pointing out that this chess thing is a good datum to temper their enthusiasm. What's wrong with that? Also these claims are not completely without merit, it's just that LLMs seem to excel at specific "cognitive" tasks and it's interesting to see where they fail.