6 ms·
If you give an LLM a word problem that involves the same math and change the names of the people in the word problem the LLM will likely generate different math
by nephy 2y ago
If you give an LLM a word problem that involves the same math and change the names of the people in the word problem the LLM will likely generate different mathematical results. Without any knowledge of how any of this works, that seems pretty damning of the fact that LLMs do not reason. They are predictive text models. That’s it.
- alexwebb2 2y agoDemonstrably false. https://chatgpt.com/share/6722ca8a-6c80-800d-89b9-be40874c5b65 https://chatgpt.com/share/6722ca8a-6c80-800d-89b9-be40874c5b... https://chatgpt.com/share/6722ca97-4974-800d-99c2-bb58c60ea632 https://chatgpt.com/share/6722ca97-4974-800d-99c2-bb58c60ea6...
- TZubiri 2y agoIt's worth noting that this may not be result of a pure LLM, it's possible that ChatGPT is using "actions", explicitly: 1- running the query through a classifier to figure out if the question involves numbers or math 2- Extract the function and the operands 3- Do the math operation with standard non-LLM mechanisms 4- feed back the solution to the LLM 5- Concatenate the math answer with the LLM answer with string substitution. So in a strict sense this is not very representative of the logical capabilities of an LLM.
- deleted 2y ago[deleted]
- thomashop 2y agoIt shows you when it's calling functions. I also did the same test with Llama, which runs locally and cannot access function calls and it works.
- TZubiri 2y agoYou are right I actually downloaded Llama to do more detailed tests. God bless Stallman.
- digging 2y agoThen what's the point of ever talking about LLM capabilities again? We've already hooked them up to other tools. This confusion was introduced at the top of the thread. If the argument is "LLMs plus tooling can't do X," the argument is wrong. If the argument is "LLMs alone can't do X," the argument is worthless. In fact, if the argument is that binary at all, it's a bad argument and we should laugh it out of the room; the idea that a lay person uninvolved with LLM research or development could make such an assertion is absurd.
- TaylorAlexander 2y agoAt this point I really only take rigorous research papers in to account when considering this stuff. Apple published research just this month that the parent post is referring to. A systematic study is far more compelling than an anecdote. https://machinelearning.apple.com/research/gsm-symbolic https://machinelearning.apple.com/research/gsm-symbolic
- famouswaffles 2y agoThat study shows 4o, o1-mini and o1-preview's new scores are all within margin error on 4/5 of their new benchmarks(some even see increases). The one that isn't involves changing more than names. Changing names does not affect the performance of Sota models.
- gruez 2y ago>That study very clearly shows 4o, o1-mini and o1-preview's new scores are all within margin error on 4/5 of their new benchmarks. Which figure are you referring to? For instance figure 8a shows a -32.0% accuracy drop when an insignificant change was added to the question. It's unclear how that's "within the margin of error" or "Changing names does not affect the performance of Sota models".
- famouswaffles 2y agoTable 1 in the Appendix. GSM-No-op is the one benchmark that sees significant drops for those 4 models as well (with preview dropping the least at -17%). No-op adds "seemingly relevant but ultimately inconsequential statements". So "change names, performance drops" is decidedly false for today's state of the art.
- gruez 2y agoThanks. I wrongly focused on the headline result of the paper rather than the specific claim in the comment chain about "changing name, different results".
- 2y ago
- gruez 2y agoTo be fair, the claim wasn't that it always produced the wrong answer, just that there exists circumstances where it does. A pair of examples where it was correct hardly justifies a "demonstrably false" response.
- thomashop 2y agoConversely, a pair of examples where it was incorrect hardly justifies the opposite response. If you want a more scientific answer there is this recent paper: https://machinelearning.apple.com/research/gsm-symbolic https://machinelearning.apple.com/research/gsm-symbolic
- EraYaN 2y agoIt kind of does though, because it means you can never trust the output to be correct. The error is a much bigger deal than it being correct in a specific case.
- thomashop 2y agoYou can never trust the outputs of humans to be correct but we find ways of verifying and correcting mistakes. The same extra layer is needed for LLMs.
- digging 2y ago> It kind of does though, because it means you can never trust the output to be correct. Maybe some HN commenters will finally learn the value of uncertainty then.
- astrange 2y agoMinor edits to well known problems do easily fool current models though. Here's one 4o and o1-mini fail on, but o1-preview passes. (It's the mother/surgeon riddle so kinda gore-y.) https://chatgpt.com/share/6723477e-6e38-8000-8b7e-73a3abb652a9 https://chatgpt.com/share/6723477e-6e38-8000-8b7e-73a3abb652... https://chatgpt.com/share/6723478c-1e08-8000-adda-3a378029b465 https://chatgpt.com/share/6723478c-1e08-8000-adda-3a378029b4... https://chatgpt.com/share/67234772-0ebc-8000-a54a-b597be3a1f4f https://chatgpt.com/share/67234772-0ebc-8000-a54a-b597be3a1f...
- Workaccount2 2y agoThis is a relatively trivial task for current top models. More challenging are unconventional story structures, like a mom named Matthew with a son named Mary and a daughter named William, who is Matthew's daughter? But even these can still be done by the best models. And it is very unlikely there is much if any training data that's like this.
- alexwebb2 2y agoThat's a neat example problem, thanks for sharing! For anyone curious: https://chatgpt.com/share/6722d130-8ce4-800d-bf7e-c1891dfdf781 https://chatgpt.com/share/6722d130-8ce4-800d-bf7e-c1891dfdf7... > Based on traditional naming conventions, it seems that the names might have been switched in this scenario. However, based purely on your setup: > > Matthew has a daughter named William and a son named Mary. > > So, Matthew's daughter is William.
- rileymat2 2y agoHow do people fair on unconventional structures? I am reminded of that old riddle involving a the mother being the doctor after a car crash.
- adwn 2y agoNo idea why you've been downvoted, because that's a relevant and true comment. A more complex example would be the Monty Hall problem [1], for which even some very intelligent people will intuitively give the wrong answer, whereas symbolic reasoning (or Monte Carlo simulations) leads to the right conclusion. [1] https://en.wikipedia.org/wiki/Monty_Hall_problem https://en.wikipedia.org/wiki/Monty_Hall_problem
- jklinger410 2y agoThis is what kind of comments you make when your experience with LLMs is through memes.
- vanviegen 2y agoAnd yet, humans, our benchmark for AGI, suffer from similar problems, with our reasoning being heavily influenced by things that should have been unrelated. https://en.m.wikipedia.org/wiki/Priming_(psychology) https://en.m.wikipedia.org/wiki/Priming_(psychology)