4 ms·
This (GPT-3’s performance on arithmetic tasks) is covered in the original paper (https://arxiv.org/abs/2005.14165 https://arxiv.org/abs/2005.14165) Pages: 21-2
by jointpdf 6y ago
This (GPT-3’s performance on arithmetic tasks) is covered in the original paper (https://arxiv.org/abs/2005.14165 https://arxiv.org/abs/2005.14165)
Pages: 21-23, 63
- YeGoblynQueenne 6y agoSee my reply above. The evidence in the paper does not support the claim that GPT-3 has learned rules of arithmetic (also claimed in the paper, not just in the OP's comment). There is a much simpler explanation, i.e. memorisation.
- jointpdf 6y agoYour reply is misleading, sorry. You didn’t offer any actual evidence of your own to support the memorization claim. You didn’t even do your own arithmetic problem correctly. Since your performance on this task was <100% accurate, I can only assume you do not know the rules of arithmetic. > How many "x + y" questions can be formulated where x and y are both single-digit numbers? The answer is 2^10, or 100. Less snarkily, if there’s (10^4)^2 = 100 million combinations of 4 digit addition problems, and GPT-3 is reaching 25.5% accuracy on those problems (vs. 0.4% in the 13B parameter model). For 3 digit problems, it’s even better: 1 million combinations and 80.4% accuracy. Clearly, there is more happening than simple memorization—the training set does not contain 800k 3-digit addition problems. Thus, it’s fair to say that the model has at least a partial grasp of how to perform arithmetic operations (but probably not fair to say that it has synthesized the entire system of arithmetic). Also, the paper does say that they scrubbed exact examples from the training set to avoid memorization, a fact you left out: > (pg. 23): ”To spot-check whether the model is simply memorizing specific arithmetic problems, we took the 3-digit arithmetic problems in our test set and searched for them in our training data in both the forms "<NUM1> + <NUM2> =" and "<NUM1> plus <NUM2>". Out of 2,000 addition problems we found only 17 matches (0.8%) and out of 2,000 subtraction problems we found only 2 matches (0.1%), suggesting that only a trivial fraction of the correct answers could have been memorized. In addition, inspection of incorrect answers reveals that the model often makes mistakes such as not carrying a “1”, suggesting it is actually attempting to perform the relevant computation rather than memorizing a table.”
- fossuser 6y ago> "In addition, inspection of incorrect answers reveals that the model often makes mistakes such as not carrying a “1”, suggesting it is actually attempting to perform the relevant computation rather than memorizing a table.”" This is super interesting and something I hadn't read before. That is very cool, and definitely suggests it's figuring out how the computation actually works (!).
- YeGoblynQueenne 6y agoThis is a very low standard of evidence. The model answers an arithmetic problem correctly - "It has learned arithmetic!". The model answers an arithmetic problem incorrectly - "it has learned arithmetic!". What sense does that make?
- fossuser 6y agoIt's because it's not binary. The nature of the failure (errors where it 'forgot' to carry the one) suggest that it's doing something like basic arithmetic and making mistakes. This is evidence in the direction of having a model of how to do basic arithmetic and evidence against memorization. I'm not pretending that both outcomes mean it knows arithmetic. For example, if the outputs were random or if they only matched exact examples it had seen then it would look like memorization, but that isn't what's seen.
- YeGoblynQueenne 6y agoAs I say in my previous comment the "evidence" is of a very low standard. This is what's reported in the paper: In addition, inspection of incorrect answers reveals that the model often makes mistakes such as not carrying a “1”, suggesting it is actually attempting to perform the relevant computation rather thanmemorizing a table. So, what is "often"? 100% of the time? 60% of the time? 30% of the time? Such a vague statement is no evidence of anything, much less the very strong claim made in the paper.
- YeGoblynQueenne 6y ago
- typon 6y agoThat's right. I don't even know how to define "learning the rules of arithmetic" in the context of a neural network.