3 ms·
I’m not attacking you, which is why I immediately offered you a salve for the burn by marking my comment as snarky. It did have a real point though—that learnin
by jointpdf 6y ago
I’m not attacking you, which is why I immediately offered you a salve for the burn by marking my comment as snarky. It did have a real point though—that learning (even something inherently logical like arithmetic) is not a binary outcome, as fossuser also pointed out.
The claim that “GPT-3’s performance on arithmetic tasks is solely due to memorization / data leakage—it has no generalization ability on this type of task”, is easily attackable by...well, doing arithmetic and applying common sense.
There are 2(10^5)^2 = 20 billion possible 5 digit problems (both addition and subtraction). The accuracy on those tasks is about 10%, so roughly 2 billion* 5-digit addition and subtraction problems would need to be represented in the training data (Common Crawl + books as you said). Each problem is at minimum 5 tokens (e.g. 99999 + 11111 = 111110). So is ~2.5% of the training corpus 5-digit addition and subtraction problems that eluded their filtration process? (assuming it’s ~400B tokens like you said). Seems exceedingly unlikely, so much so that memorization ceases to be the simplest explanation.
That said, yes it is surprising that a language model can generalize in this way—that’s the point of the paper. How exactly this happens seems like a valuable thread to pull. Your critiques may help, but writing the results off as impossible magic does not.
- YeGoblynQueenne 6y agoI was a bit touchy yesterday, I guess - thanks for the salve. About the 5-digit problems- I didn't make this clear but I don't think those were memorised. I think the two- and three-digit problems (all three operations) were memorised, because those are the most likely to be represented in their entirety, or close, in GPT-3's training corpus, given that they are operations that are common to very common in daily life. I doubt that the four- and five-digit addition problems (and the single-digit, multi-op problem) were represented often enough in GPT-3's training corpus for them to be memorised. I think the low accuracy in these problems (less than 10% in the few-shot setting and near zero in the zero- and one-shot) is low enough that it doesn't require an explanation other than a mix of luck and overfitting that is common enough in machine learning algorithms that it's no surprise. e.g. we evaluate classifiers using diverse metrics, not just accuracy, because this is so common. It is this observation, that GPT-3 did well in problems that are likely to be well reprsented in its training corpus and badly in ones that aren't, that convinces me that no more complicated explanation is needed than memorisation. Something else. Like I say above, we evaluate classifiers not only by accuracy (the rate of correct answers), because accuracy can be misleading. e.g. a classifier can have 100% accuracy with 0% false positives and 100% false negatives. The GPT-3 authors only tested the ability of their model to give answers to problems stated as "x + y = ". They didn't test, e.g. what happens if they prompt it with "10 + 20 = 40, 38 + 25 = ". Testing for aberrant answers following from such confusing prompts has often showed that language models that appear to be answering questions correctly because of a deep understanding of language are in truth overfitting to surface statistical regularities. See for example [1,2] and many other references in [3]. Indeed, I could be wrong about rote memorisation and GPT-3 can still not be learning to perform arithmetic computations, given the tendency of language models to learn spurious correlations. There is an article about a mathemagician on the front page today, that shows how she found roots of huge numbers by finding shortcuts around expensive calculations. For instance, all sums between numbers ending in 5 end in 0, etc. I wouldn't find it magickal if a language model was finding such heuristics and that this is the "something else" that is said to be going on. However that would not be "generalisation" and it would not be learning to perform arithmetic. In the end, I don't understand how a model can be said to know how to add two-digit numbers perfectly but not five-digit numbers. If it's performing an incomplete computation in the latter case, then what kind of incomplete computation is it performing? If it "gets it wrong after three digits" then why does it get three digits right? What's the big difference between three- and four-digit numbers that causes performance to fall off a cliff - other than the chance of finding such numbers in a natural language corpus? As to magick- I'm writing off not the results, but the hand-waving presented in place of an explanation as magick. GPT-3 is a technological artifact designed to do one job, now reported to be doing another. This requires a thorough explanation but instead we got magickal thinking: the authors wish that their models could learn arithmetic, so they took its behaviour as proof that it learned arithmetic. ___________________ [1] Probing Neural Network Comprehension of Natural Language Arguments https://www.aclweb.org/anthology/P19-1459/ https://www.aclweb.org/anthology/P19-1459/ [2] Right for the Wrong Reasons: Diagnosing Syntactic Heuristics in Natural Language Inference https://www.aclweb.org/anthology/P19-1334/ https://www.aclweb.org/anthology/P19-1334/ [3] https://www.technologyreview.com/2020/07/31/1005876/natural-language-processing-evaluation-ai-opinion/ https://www.technologyreview.com/2020/07/31/1005876/natural-... (try F9 if you 're over limit)
- jointpdf 6y agoI did feel genuinely bad for my comment, since it was a few shades nastier than I was going for and I try hard to be positive and not cut people down. It was mean spirited and I apologize. My point was that intelligent humans can and often do make mistakes in logic and computation (arithmetic) in ways that machines typically do not. One reason may be colliding or incomplete representations of certain concepts, and (relatedly) the fact that we are relying on language. I think of neural networks as fuzzy representation composers, so it seems they also fail for similar reasons. Basically, it (GPT) does have some layered representation of the concept of numbers and how they are used in different contexts which gives it some faculty at carrying out common operations, but it doesn’t “add up” to a reliable system of logic (that would allow it to extend addition to say 100-digit numbers, the way even a sharp and/or patient 2nd grader could do, generalizing from the simpler cases). I think accuracy is sometimes the correct measure, and in this instance it seems fine—at baseline, we should expect ~0% accuracy since it is generating output from essentially the space of all possible text (texts <= 2048 tokens). I agree that it would be interesting to probe the model with better tests, and understanding when/why it fails on certain arithmetic problems or types of reasoning. I liked what you wrote about finding heuristics, though I disagree with your conclusion that heuristic finding does not qualify as learning—it is just somewhere along the spectrum between a randomized model and an ALU (neither of which can be said to have learned anything) in terms of its ability to perform arithmetic. Of course, we already have better models for solving proofs and such, so I generally think the way toward more complete AI models is to return to the system design view of AI (meta-learning, integration of different models, etc) rather than trying to evolve one colossal model to rule them all. That is, a meta-model that recognizes what sort of problem it is facing, then selecting a model/program to solve or generate possible solutions to that problem, while revealing or explaining as much of this process as possible to the user. In any case, I have definitely have more to read on the subject and am mostly musing at this point. Thanks for the references and the conversation.
- YeGoblynQueenne 6y agoYou're welcome, and really, please don't worry about your comment. It's all good :) I guess I can concede that the memorisation explanation is not the only possible one, there's always the possibility of learned heuristics. I still expect very strong evidence before I'm convinced that GPT-3 can learn arithmetic in the general sense and I don't trust the explanation that it's only learning partially- but let's agree to disagree on that. Thank you for the conversation, too.