3 ms·
I mean obviously yes. LLMs aren't intelligent and they don't understand anything. OP is trying to argue against this but it's nonsense. ChatGPT also cannot tel
by user_named 3y ago
I mean obviously yes. LLMs aren't intelligent and they don't understand anything.
OP is trying to argue against this but it's nonsense. ChatGPT also cannot tell me who the 8th chancellor was. It tells me it was Gerhard Schröder.
- user_named 3y agoIt also cannot tell me the 7th.
- TeMPOraL 3y agoIDK, I still some of it boil down to "B -> A" being extremely ambiguous. Other than the annoying Tom Cruise example (I finally know where it came from!), I've seen people frequently bring up "The color of the sky is blue" vs. "Blue is the color of the sky" - but this illustrates my hypothesis perfectly. In training data (and in the totality of what humans ever said or wrote), "Color of the sky is " is almost certain to be followed by "blue" - meanwhile, "Blue is the color of the " has a lot of possible completions, of which "sky" is by far not the most likely.
- user_named 3y agoNah. The context of all of this is people claiming LLMs are AI. If so they would be able to reverse A is B, and also know when not to, so your example is irrelevant. The paper shows that it cannot reverse. Then the blog posts goes around different disingenuous ways of making excuses for why it can't reverse at the same time as trying to show that it can reverse in the Tom Cruise training example (extremely disingenuous because it's just the same data over and over with synonyms replaced, and what he gets out of it is just completions, not logical deduction).
- TeMPOraL 3y agoI still call bullshit on this. Humans don't do reversals for free either. We may memorize the reversal, or first, infer it. That's an extra step, often very cognitively challenging. I'd like to see experiments showing how "reversal curse" fares when LLM is allowed to make the inference. Like: "Please complete the sentence: [your B->A] - but reason through the process first. Start with identifying the subject, then write a high-level summary of what you know about the subject, and only then attempt to complete the original sentence." I'd expect something like this to suddenly score much better. And I don't consider this cheating - because I think the "AI-ness" quality of LLMs shouldn't be measured against the workings of a human mind, but rather the workings of the inner voice in the human's mind.
- foldr 3y agoThis simply isn't true of humans for the kinds of examples used in the paper. If I read "Olaf Scholz was the ninth Chancellor of Germany" then I have no trouble reversing that and determining that the 9th chancellor of Germany was Olaf Scholz. This is not an inference that's 'cognitively challenging'. Humans may sometimes fail to make simple inferences of this sort when building their factual databases, but they don't systematically fail to do so. Also, you may have missed this part of the paper: >The Reversal Curse shows a basic inability to generalize beyond the training data. Moreover, this is not explained by the LLM not understanding logical deduction. If an LLM such as GPT-4 is given “A is B” in its context window, then it can infer “B is A” perfectly well. The paper does not claim that GPT-4 cannot perform logical deduction, but only that it does not appear to make use of it when generalizing its training data.
- TeMPOraL 3y ago> This simply isn't true of humans for the kinds of examples used in the paper. If I read "Olaf Scholz was the ninth Chancellor of Germany" then I have no trouble reversing that and determining that the 9th chancellor of Germany was Olaf Scholz. This is not an inference that's 'cognitively challenging'. That's in-context though. LLMs don't fail in-context either. For a better comparison, recall your school experience, say with history lessons, or geography lessons - where you would cram a hundred "A is B" relationships, and then take a test that demanded you know the reversals. Not as easy. The "reversal curse" failures I've seen with LLMs are very similar to asking a random person, out of a blue, some unusual reversal of some random fact they ought to know, and then being surprised they can't answer quickly. > The paper does not claim that GPT-4 cannot perform logical deduction, but only that it does not appear to make use of it when generalizing its training data. Well, neither can humans when cramming, if you don't give them time to pause and think about what they're learning. I believe the equivalent is happening here - LLMs can perform logical deductions, but at no point in the training process is this capability used.
- foldr 3y ago>For a better comparison, recall your school experience, say with history lessons, or geography lessons - where you would cram a hundred "A is B" relationships, and then take a test that demanded you know the reversals. Not as easy. This doesn't correspond to my school experience. In general I don't feel that I have to separately memorise "A is B" and "B is A". For example, if I learn that Elizabeth I was Henry VIII's daughter, I don't also have to learn that he was her father. >LLMs can perform logical deductions, but at no point in the training process is this capability used. That is exactly what the paper says.