3 ms·
It's hard to define "reason". Certainly, the pattern matching is good enough to solve certain equations step by step, for example, and that's one form of reason
by codeflo 2y ago
It's hard to define "reason". Certainly, the pattern matching is good enough to solve certain equations step by step, for example, and that's one form of reasoning.
But just ask ChatGPT (4o) how a farmer and a sheep would cross a river on a boat. For certain formulations, it reliably hallucinates a wolf or a piece of cabbage into the problem. Why? Because it pattern matches the solution to a well-known puzzle.
We've now seen enough failure modes of these models to realize that a lot of the early successes in logic tests were us not realizing how deep the corpus actually is.
Basically any question humans have ever asked or thought of is in the training data. And ChatGPT has remembered so much of it that it's very hard to surprise the model. But when we manage to do, we see it's actually dumber than a three-year-old. (Which to be fair, is actually still an accomplishment that just a few years ago almost nobody would have considered possible.)
- EnergyAmy 2y agoIMO a fair evaluation needs to include a few extra steps, where the model generates a response, and then reflects on that response, using basic prompts like "Did you make sure to read the question carefully and look for trick questions?". Imagine if you were tested in a rapid-fire manner and evaluated based on whatever answer you blurted out first. Human reasoning scores would suffer dramatically as well. There would probably be a lot of answers involving wolves and cabbages. You might think that prompting the LLM with questions like "Make sure it's not a trick question" is cheating, but it's very similar to how humans work. A lot of people would answer wrongly if they encountered the question "What weighs more, a pound of bricks or two pounds of feathers?", because they also just pattern match and assume the answer is the question they've seen many times before. If they're primed by someone saying "Read the question carefully", they'll do a lot better.
- Hugsun 2y agoYou might enjoy the analysis in the article. https://www.arnaldur.be/writing/about/large-language-model-reasoning#recognizability-traps https://www.arnaldur.be/writing/about/large-language-model-r... There is an expandable box with all the questions, colored by correctness, and question 6 has a bunch of responses. Below the box is a summary of the results.
- Hugsun 2y ago> But just ask ChatGPT (4o) how a farmer and a sheep would cross a river on a boat. For certain formulations, it reliably hallucinates a wolf or a piece of cabbage into the problem. Why? Because it pattern matches the solution to a well-known puzzle. I analyze that exact anti-puzzle in the chapter Recognizability traps in the post. https://www.arnaldur.be/writing/about/large-language-model-reasoning#recognizability-traps https://www.arnaldur.be/writing/about/large-language-model-r... In case you didn't notice, there is an expandable box with the results.
- EnergyAmy 2y agoI'd be interested in seeing an evaluation with a simple follow-up question of something like "Is that the best solution?" or "Did you check for trick questions?". I'm able to replicate the error with the original question, but asking "Is that the best solution?" makes it recognize its error and fix it. Also, I've gotten better results from GPT-4 than GPT-4o for purely text-based tasks. I rather wonder if OpenAI is pushing GPT-4o primarily because it's cheaper compute for them or something.
- Hugsun 2y agoLook at question six and its responses. That is what I look into there. From the article: > I asked the model further about its 6th response. It realized its error at the slightest hint of disagreement, but just [asking it to elaborate] didn’t help; it was reliably incorrect.
- EnergyAmy 2y agoYeah, you asked it specific follow-up questions, designed to guide it towards the correct answer. That feels like cheating, because it's not generalizable. I'd like to see an evaluation with a generic prompt like "Is that the best solution? Make sure you look for trick questions.", because that is applicable to any input. Or something like "Pick apart this answer like a pedantic HN commenter. Do your best to prove it wrong, or begrudgingly say it's correct if you can't: <previous answer here>"