5 ms·
All the environments the test (Tower of Hanoi, Checkers Jumping, River Crossing, Block World) could easily be solved perfectly by any of the LLMs if the authors
by thomasahle 1y ago
All the environments the test (Tower of Hanoi, Checkers Jumping, River Crossing, Block World) could easily be solved perfectly by any of the LLMs if the authors had allowed it to write code.
I don't really see how this is different from "LLMs can't multiply 20 digit numbers"--which btw, most humans can't either. I tried it once (using pen and paper) and consistently made errors somewhere.
- someothherguyy 1y ago> humans can't The reasons humans can't and the reasons LLMs can't are completely different though. LLMs are often incapable of performing multiplication. Many humans just wouldn't care to do it.
- Jensson 1y ago> I don't really see how this is different from "LLMs can't multiply 20 digit numbers"--which btw, most humans can't either. I tried it once (using pen and paper) and consistently made errors somewhere. People made missiles and precise engineering like jet aircraft before we had computers, humans can do all of those things reliably just by spending more time thinking about it, inventing better strategies and using more paper. Our brains weren't made to do such computations, but a general intelligence can solve the problem anyway by using what it has in a smart way.
- thomasahle 1y agoSome specialized people could probably do 20x20, but I'd still expect them to make a mistake at 100x100. The level we needed for space crafts was much less than that, and we had many levels of checks to help catch errors afterwards. I'd wager that 95% of humans wouldn't be able to do 10x10 multiplication without errors, even if we paid them $100 to get it right. There's a reason we had to invent lots of machines to help us. It would be an interesting social studies paper to try and recreate some "LLMs can't think" papers with humans.
- Jensson 1y ago> There's a reason we had to invent lots of machines to help us. The reason was efficiency, not that we couldn't do it. If a machine can do it then we don't need expensive humans to do it, so human time can be used more effectively.
- moralestapia 1y agoI don't think you got @Jensson's point. With enough effort and time we can arrive at a perfect solution to those problems without a computer. This is not a hypothetical, it was like that for at least hundreds of years.
- throw310822 1y agoWith enough time and effort you can build an entire science of how arbitrarily complex computations can be done with pen and paper without errors in an arbitrarily long amount of time. But then you're not measuring the ability to perform the calculations, but the ability to invent the methods that make the calculation possible.
- jdmoreira 1y agoNo. a huge population of humans did while standing on the shoulders of giants.
- Jensson 1y agoHumans aren't giants, they stood on the shoulder of other humans. So for AI to be equivalent they should stand on the shoulders of other AI models.
- jdmoreira 1y agobuilding for thousands of years with a population size in the range between millions and billions at any given time.
- Jensson 1y agoRight, and when we have AI that can do the same with millions/billions of computers then we can replace humans. But as long as AI cannot do that they cannot replace humans, and we are very far from that. Currently AI cannot even replace individual humans in most white collar jobs, and replacing entire team is way harder than replacing an individual, and then even harder is replacing workers in an entire field meaning the AI has to make research and advances on its own etc. So like, we are still very far from AI completely being able to replace human thinking and thus be called AGI. Or in other words, AI has to replace those giants to be able to replace humanity, since those giants are humans.
- Xmd5a 1y ago>Large Language Model as a Policy Teacher for Training Reinforcement Learning Agents >In this paper, we introduce a novel framework that addresses these challenges by training a smaller, specialized student RL agent using instructions from an LLM-based teacher agent. By incorporating the guidance from the teacher agent, the student agent can distill the prior knowledge of the LLM into its own model. Consequently, the student agent can be trained with significantly less data. Moreover, through further training with environment feedback, the student agent surpasses the capabilities of its teacher for completing the target task. https://arxiv.org/abs/2311.13373 https://arxiv.org/abs/2311.13373
- hskalin 1y agoWell that's because all these LLMs have memorized a ton of code bases with solutions to all these problems.
- bwfan123 1y ago> but humans cant do it either This argument is tired as it keeps getting repeated for any flaws seen in LLMs. And the other tired argument is: wait ! this is a sigmoid curve, and we have not seen the inflection point yet. If someone have me a penny for every comment saying these, I'd be rich by now. Humans invented machines because they could not do certain things. All the way from simple machines in physics (Archimedes lever) to the modern computer.
- thomasahle 1y ago> Humans invented machines because they could not do certain things. If your disappointment is that the LLM didn't invent a computer to solve the problem, maybe you need to give it access to physical tools, robots, labs etc.
- mrbungie 1y agoNah, even if we follow such a weak "argument" the fact is that, ironically, the evidence shown in this and other papers point towards the idea that even if LRMs did have access to physical tools, robots labs, etc*, they probably would not be able to harness them properly. So even if we had an API-first world (i.e. every object and subject in the world can be mediated via a MCP server), they wouldn't be able to perform as well as we hope. Sure, humans may fail doing a 20 digit multiplication problems but I don't think that's relevant. Most aligned, educated and well incentivized humans (such as the ones building and handling labs) will follow complex and probably ill-defined instructions correctly and predictably, instructions harder to follow and interpret than an exact Towers of Hanoi solving algorithm. Don't misinterpret me, human errors do happen in those contexts because, well, we're talking about humans, but not as catastrophically as the errors committed by LRMs in this paper. I'm kind of tired of people comparing humans to machines in such simple and dishonest ways. Such thoughts pollute the AI field. *In this case for some of the problems the LRMs were given an exact algorithm to follow, and they didn't. I wouldn't keep my hopes up for an LRM handling a full physical laboratory/factory.
- thomasahle 1y ago
- mjburgess 1y agoThe goal isnt to assess the LLM capability at solving any of those problems. The point isnt how good they are at block world puzzles. The point is to construct non-circular ways of quantifying model performance in reasoning. That the LLM has access to prior exemplars of any given problem is exactly the issue in establishing performance in reasoning, over historical synthesis.
- thomasahle 1y agoHow are these problems more interesting than simple arithmetic or algorithmic problems?
- mrbungie 1y agoTowers of Hanoi IS an algorithmic problem. It is a high-school/college level problem when designing algorithms, probably kid level when trying to solve intuitively, heuristically or via brute force for few disks (i.e. like when playing Mass Effect 1 or similar games that embed it as a minigame*). * https://www.youtube.com/watch?v=1vTBVyhX7n4 https://www.youtube.com/watch?v=1vTBVyhX7n4
- pcooperchi 1y agoThe problems themselves aren’t particularly interesting, I suppose. The interesting part is how the complexity of each problem scales as a function of the number of inputs (e.g. the number of disks in the tower of Hanoi).
- blither 1y ago> if the authors had allowed it to write code. Yeah, and FWIW doing this through writing code is trivial in an LLM / LRM - after testing locally took not even a minute to have a working solution no matter the amount of disks. Your analogy makes sense, no reasonable person would try to solve a Tower of Hanoi type problem with e.g. 15 disks and sit there for 32,767 moves non-programmatically.
- PatronBernard 1y ago>write code Doesn't that come down to allowing it to directly regurgitate training data? Surely it's seen dozens of such solutions.