5 ms·
Reasoning is meaningfully distinct from what the models do now. You can ask ChatGPT to multiply two smallish natural numbers and often it will give the correct
by ookdatnog 4y ago
Reasoning is meaningfully distinct from what the models do now. You can ask ChatGPT to multiply two smallish natural numbers and often it will give the correct answer. Now look what happens if you ask it to multiply two very large numbers:
> To multiply two large numbers, such as 3146746135164684 times 468473136468433, we can use the traditional multiplication algorithm. This algorithm involves breaking the numbers down into smaller groups of digits, multiplying each group by the corresponding group of digits in the other number, and then adding up the products to find the final result.
> Here is an example of how to multiply 3146746135164684 times 468473136468433 using the traditional multiplication algorithm:
> Break the numbers down into groups of digits. For example, we could group the digits of 3146746135164684 as 314, 6746, 135, 1646, 84 and group the digits of 468473136468433 as 4684, 7313, 6468, 433.
> Multipply each group of digits by the corresponding group of digits in the other number. For example, we would multiply 314 times 4684 to get 1470196, 6746 times 7313 to get 49470158, 135 times 6468 to get 873980, 1646 times 433 to get 708338, and 84 times 1 to get 84.
> Add up the products to find the final result. In this case, the final result would be 1470196 + 49470158 + 873980 + 708338 + 84 = 146845556.
> Therefore, the result of multiplying 3146746135164684 times 468473136468433 using the traditional multiplication algorithm is 146845556.
It's not just that the answer is wrong, is that it's complete nonsense.
Reasoning is a style of thinking that scales. You may be more likely to get the wrong answer in a very long chain of reasoning because at every step you have a nonzero chance of making a mistake, but the mistake is identifiable and explainable. That's why teachers ask you to show your work. Even if you get the answer wrong, they can see at a glance whether you understand the material or not. We can see at a glance that ChatGPT does not understand multiplication.
- johnfn 4y agoI don't think I buy this argument. ChatGPT seems to understand how to reason about a large multiplication the same that a 6 or 7 year old might, and I would expect a 6 or 7 year old to make similarly large errors. No one claims that 6 or 7 year olds are unable to reason.
- gonehome 4y agoYeah, in the original gpt-3 paper one of the more interesting bits was that it made similar off by one errors a human would make when doing arithmetic (and they controlled for memorized test data).
- idontpost 4y ago> ChatGPT seems to understand how to reason about a large multiplication the same that a 6 or 7 year old might This is purely ignorant magical thinking. You're ignoring the very real facts of how ChatGPT works and fabricating a fantasy that you prefer to believe.
- johnfn 4y agoHardly. Go ask a 6 or 7 year old to write down their "reasoning" at multiplying 203948029384029384 by 7834928734982374982374 and, assuming you can even get them to do it, I argue that the words they write will be very similar to what ChatGPT wrote. Perhaps the vocabulary will be less refined, but the essence will be similar. In any case, the point is that it's an objective thing anyone can observe: compare the two outputs. See if they differ. > This is purely ignorant magical thinking. I counter that you ascribe "magical" thinking to the human condition.
- imperfect_blue 4y agoA 7 year old would know that you would want to use a calculator, and indeed, would be capable of multiplying those numbers with a calculator.
- ben_w 4y agoBut by that standard, I could just ask GPT to write code to multiply the numbers. It's not even trying you be an embodied agent with a hand able to type onto a calculator. GPT is remarkable to me not for its failings, but because of how far it can get when it's only trying to win the game of "predict what token comes next".
- sdenton4 4y agoSub-token manipulation is a known weakness of the models; they are fed blocks of characters, not individual characters, and so they underperform on tasks that require fine-character-level manipulations.
- kqr 4y agoBut in that case I would have wanted to see it present a number that's at least of the right order of magnitude, even if all individual digits were incorrect. A human can adapt to their known weaknesses.
- sdenton4 4y agoOptical illusions expose fundamental issues in human perception, which we generally can't adapt to. This is similar. [edit, fwiw]: Also consider that humans can't perceive ultrasonic sound frequencies, or light outside a certain frequency range... We have eventually invented ways to perceive beyond our physical limitations, but no amount of individual adaptation can overcome these boundaries. There are /other/ arguments that ChatGPT doesn't reason well, but the character-level manipulation examples are insufficient.
- zarzavat 4y agoThis post would be more compelling if you multiplied those numbers out yourself to prove that a human can do it. I myself don't feel comfortable to multiply those numbers. I'd say your definition of reasoning excludes most humans, who would simply respond to such a prompt with "nope, can't do it".
- ookdatnog 4y agoI believe there is a meaningful difference between an answer that's wrong and an answer that's nonsense. If I were to multiply those numbers, it's likely that I'd get the wrong outcome because it's a long computation and at every step I have a nonzero chance of making a mistake. My solution, written out, would look like the correct algorithm, but with computational mistakes along the way. My result could be quite far off, but it would be in roughly the same order of magnitude. If I'd get an answer that's not roughly in the order of magnitude that I expect, I would spot it and -- if so motivated -- start over. If you'd look at my work, you would be able to conclude that I understand multiplication but made a computational error. ChatGPT also diligently describes its work, and it's just nonsense. The final result is smaller than either of the factors. The algorithm it uses makes no sense. On smaller numbers, the algorithm also doesn't make sense, but it can ballpark the outcome. Therefore, it seems to be mostly using estimation rather than computation to get to the answer, and that estimation breaks down completely for very large numbers.
- josephcsible 4y ago> I myself don't feel comfortable to multiply those numbers. Keep in mind that being super confident in a completely wrong answer is way worse than that is.
- ben_w 4y agoSure, but also very human. While all the rocket and space nerds I follow hold Musk in high regard for everything related to SpaceX, all of the civil engineering nerds I follow think TBC is a deadly disaster waiting to happen and that hyperloop is pointless, while all the neuroscience nerds I follow think Neuralink is kinda meh.