3 ms·
> The fact that the human would have a much harder time writing a useful program is irrelevant. Why? There's a box. You give it a problem, and it comes up wi
by MattCruikshank 17d ago
> The fact that the human would have a much harder time writing a useful program is irrelevant.
Why?
There's a box.
You give it a problem, and it comes up with a solution.
Why does it matter to you if the box is strictly a LLM, or if the LLM can write code that it executes?
Even neater if the box is self-contained with a local model. You provide electricity, and it comes up with solutions. Why does it matter if it can do chess "in its head", or if it has to use scratch paper?
- lionkor 17d agoThe question is to what end? This is a benchmark task, because playing chess, or solving other well-understood problems is more of a party trick than it is useful. If you let the LLM write a chess program, which it can ONLY do because there are already so many chess programs out there, then the benchmark becomes about recall of popular program source code, not chess.
- MattCruikshank 17d agoDo you want to measure the ability of the box, or measure the ability of the box with one hand tied behind its back? More to my point, I think it's stupid to have LLMs do work that should be done by programs... programs potentially written by LLMs. I'm advising people that they should think about this distinction, themselves, when they have data and want answers.
- cbolton 17d agoNeither. As I said I want to measure cognitive abilities. Your "ability of the box" is like "economic potential" in my previous comment. If that's what you want to measure, fine. But I want a deeper understanding: what is the thing doing, how is it solving problems? I want to get a sense of its abilities that is richer than a one-dimensional scale.
- MattCruikshank 17d agoI agree that it's a fascinating to crawl inside an LLM, and also to crawl inside of a human, and try to understand the processes and limitations. Like, Phineas Gage is one of the most remarkable learning opportunities we ever had. That said, it's really weird to me when people use (and judge) LLMs one way... and won't try using them another way. Like, to judge their utility, I think we should be open to letting them write code, and use the code they produce. Otherwise, it's like judging a Chromebook without an internet connection. Like, this was one of the most dishonest ads I've ever seen: https://www.youtube.com/watch?v=gDy9AUQJ3Fg https://www.youtube.com/watch?v=gDy9AUQJ3Fg This lamp, without a working power outlet? It really doesn't do anything...
- cbolton 17d agoI completely agree.
- deleted 17d ago[deleted]