3 ms·
You have a narrow definition of reasoning. Formally and technically it is solving a symbolic reasoning task through a sequence of steps. Yes we know it's not co
by Mike_12345 3y ago
You have a narrow definition of reasoning. Formally and technically it is solving a symbolic reasoning task through a sequence of steps. Yes we know it's not conscious and not human reasoning.
> At no point does the LLM know that 5+6 = 11
Does it need to "know" that (by your narrow definition of "know") in order to reason about a word math problem?
> if asked to solve a problem in which 5+6 was an implicit component of the solution but not explicitly present in the text, it would be completely lost
Can you provide an example? What makes you believe it can't be trained to solve those too? That's just a higher abstraction over the language. Add more layers, more training, etc. Many humans cannot solve basic math word puzzles that this artificial neural network can already solve.
- PaulDavisThe1st 3y agoThey do not "solve" word puzzles. They output text that appears to be the best response to the prompt, based on their training data. If the puzzle is solvable by doing this, then they get the answer right. If the puzzle is not solvable doing that, they are unlikely to get the answer right. If I ask you to multiply two (largeish) numbers together, you will be able to do so, using an algorithm/process that you can apply to the multiplication of any two numbers, whether anyone has ever told you about those numbers before or not. LLM's cannot do this. Give them a math problem that doesn't exist in their training set and they cannot solve it. This has been demonstrated many times.
- Mike_12345 3y ago> Give them a math problem that doesn't exist in their training set and they cannot solve it. They routinely solve math problems (and other reasoning tasks) that don't exist in their training set. Examples were in that paper I linked to. This is one of the incredible emergent properties of LLMs / deep neural networks. Try it out today on GPT-4. Make up your own math problems and go for it.
- PaulDavisThe1st 3y agoPROMPT: what is 19192920 * 190101271 RESPONSE: The product of 19192920 and 190101271 is: 3653424693064240 Fail on first try.
- Mike_12345 3y agoYes shifting the goal posts and finding edge cases not well suited to LLMs, and also ignoring the chain of thought prompting. It can solve math word puzzles that are not in its training set. Yes you can find these edge cases. We know about these edge cases and that's just missing the point. There are countless examples of emergent properties in these LLMs which by definition are solving tasks outside of its training set. Its reasoning has been demonstrated on examples outside of its training set.
- PaulDavisThe1st 3y agoFirst of all, multiplying two numbers together is not "shifting the goal posts", but an absolutely basic test of any system that is claimed to able to do mathematical reasoning. I know that LLM's are not well suited for this, and that's because they cannot do arithmetic (among other things). So I tried a word puzzle that would also require simple multiplication: ------------------------------ PROMPT: i am going to cycle 1600 miles, with 234 miles on gravel roads. on paved roads i will ride at 1929288282 millimeters per second but on gravel I will ride at 0.00000000202 parsecs per second. How long will the journey take? ------------------------------- Now, I have to commend GPT on its ability to understand how you solve a problem like this, though that's not really very surprising given the huge numbers of such problems that exist in written materials. It precisely broke the problem down in a way that I suppose you could call "reasoning", but I would call "copying the formula for solving puzzles like this". And how did it do with the actual math? ---------------- 0.00000000202 parsecs per second is equivalent to 7499.6103827 miles per hour (mph), which we can calculate by converting parsecs to miles (1 parsec = 3.26 light-years = 19,173,511,840,000 miles) and dividing by the number of seconds in an hour: 0.00000000202 pc/s × 19,173,511,840,000 mi/pc ÷ 3600 s/hr = 7499.6103827 mph ---------- Utterly and completely wrong. Same issue with the millimeters per second computation. It is completely obvious why LLMs cannot do this. They cannot perform even basic arithmetic reasoning, and even more fundamentally, the ONLY capability they have is to create likely responses to prompts. For some things, this is extraordinarily (and scarily) powerful. But it is not reasoning.