4 ms·
Your reply is misleading, sorry. You didn’t offer any actual evidence of your own to support the memorization claim. You didn’t even do your own arithmetic prob
by jointpdf 6y ago
Your reply is misleading, sorry. You didn’t offer any actual evidence of your own to support the memorization claim. You didn’t even do your own arithmetic problem correctly. Since your performance on this task was <100% accurate, I can only assume you do not know the rules of arithmetic.
> How many "x + y" questions can be formulated where x and y are both single-digit numbers? The answer is 2^10, or 100.
Less snarkily, if there’s (10^4)^2 = 100 million combinations of 4 digit addition problems, and GPT-3 is reaching 25.5% accuracy on those problems (vs. 0.4% in the 13B parameter model). For 3 digit problems, it’s even better: 1 million combinations and 80.4% accuracy. Clearly, there is more happening than simple memorization—the training set does not contain 800k 3-digit addition problems. Thus, it’s fair to say that the model has at least a partial grasp of how to perform arithmetic operations (but probably not fair to say that it has synthesized the entire system of arithmetic).
Also, the paper does say that they scrubbed exact examples from the training set to avoid memorization, a fact you left out:
> (pg. 23): ”To spot-check whether the model is simply memorizing specific arithmetic problems, we took the 3-digit arithmetic problems in our test set and searched for them in our training data in both the forms "<NUM1> + <NUM2> =" and "<NUM1> plus <NUM2>". Out of 2,000 addition problems we found only 17 matches (0.8%) and out of 2,000 subtraction problems we found only 2 matches (0.1%), suggesting that only a trivial fraction of the correct answers could have been memorized. In addition, inspection of incorrect answers reveals that the model often makes mistakes such as not carrying a “1”, suggesting it is actually attempting to perform the relevant computation rather than
memorizing a table.”
- fossuser 6y ago> "In addition, inspection of incorrect answers reveals that the model often makes mistakes such as not carrying a “1”, suggesting it is actually attempting to perform the relevant computation rather than memorizing a table.”" This is super interesting and something I hadn't read before. That is very cool, and definitely suggests it's figuring out how the computation actually works (!).
- YeGoblynQueenne 6y agoThis is a very low standard of evidence. The model answers an arithmetic problem correctly - "It has learned arithmetic!". The model answers an arithmetic problem incorrectly - "it has learned arithmetic!". What sense does that make?
- fossuser 6y agoIt's because it's not binary. The nature of the failure (errors where it 'forgot' to carry the one) suggest that it's doing something like basic arithmetic and making mistakes. This is evidence in the direction of having a model of how to do basic arithmetic and evidence against memorization. I'm not pretending that both outcomes mean it knows arithmetic. For example, if the outputs were random or if they only matched exact examples it had seen then it would look like memorization, but that isn't what's seen.
- YeGoblynQueenne 6y agoAs I say in my previous comment the "evidence" is of a very low standard. This is what's reported in the paper: In addition, inspection of incorrect answers reveals that the model often makes mistakes such as not carrying a “1”, suggesting it is actually attempting to perform the relevant computation rather thanmemorizing a table. So, what is "often"? 100% of the time? 60% of the time? 30% of the time? Such a vague statement is no evidence of anything, much less the very strong claim made in the paper.
- YeGoblynQueenne 6y agoI don't appreciate your snark at all. I made a mistake and didn't read the paper carefully again so I confused myself with what they mean by one- two- and three-digit tasks, I accept that. I don't see what you get from pouncing on my mistake, other than a few marks for an internet burn. Now, the two- and three digit addition and subtraction tasks (operations on numbers between 0 and 99 and 0 and 999, respectively) are both small enough for the large, 175B parameter model to have memorised them exactly. Even if there was a single parameter for each three-digit number, of which there are a million, you could fit the entire set 175 thousand times in the 175 billion model (assuming they mean "a billion" as "one thousand million", not "one million million", which they don't clarify, but to be on the safe side let's assume the smallest). There is plenty of room. These four tasks are also the tasks that are most likely to be present in their entirety in a corpus of natural language, as the one GPT-3 was trained on, for example as records of common monetary transactions (especially the two-digit ones). That is, yes, the training set can comfortably contain 800k 3-digit addition problems. Why not? It contained 410 billion tokens from the Common Crawl dataset alone, plus a few extras. In short, the almost perfect accuracy on this task is not impressive. The 25% ish accuracy on the four-digit addition task is even less impressive. I don't know what the baseline is here, but 25% accuracy on anything is not something to write home about. You ask me to provide evidence of my own to support the memorisation claim. The claim is not memorisation. The claim is that GPT-3 has learned arithmetic (not stated exacly like that in the paper). This claim flies in the face of the commonly understood operation of language models, which are systems that compute the probability of a token to follow a sequnce of tokens- and nothing else. It's very hard to see how such a system should be able to perform arithmetic operations, while it's very easy to see how it can instead memorise their results. If the authors of the GPT-3 paper wish to claim that GPT-3 can perform arithmetic, instead of the much simpler explanation, they have to provide very strong evidence to back that up and refute the simpler explanation. And the "spot checks" that they performed are nowhere near such strong evidence: I can fail to find anything I search for, if I search with the wrong terms and the authors don't give much information about how they did their "spot checks". I mean, did they use a regular expression? Which one? ("<NUM1> + <NUM2> =" is not a regular expression! But then - what is it?) Did they take into account whitespace? Punctuation? Something else? What search terms they used? They dont' say. Can we tell why they failed to find what they were looking for? No. Besides, why only "spot check" three-digit arithmetic? It would make a lot more sense to spot-check two-digit problems, first, because these are the most likely to be found more often in the dataset and consequently be memorised. Indeed, the fact that they don't report "spot checks" for two-digit arithmetic suggests that they did perform those spot checks and they found a lot more overlap than for the three digit arithmetic, but chose not to report it. And if their model was memorising two-digit arithmetic, and that explains its performance on that type of task, it's safe to assume that it was memorising the third-digit arithmetic task also and that their "spot checks" were simply not very well put together to find the three-digit arithmetic examples. Note that section 4 goes in length over the possibility that the test set for all tasks (not just arithmetic) was contaminated (i.e. that it containted training examples from existing benchmarks, published on the internet). I haven't read that one carefully but test set contamination is another possibility. And, to be frank, any possibility is more possible than the possibility that a langauge model has learned arithmetic- which is tantamount to magick.