5 ms·
Gary Marcus - the author of this - has previously offered several concrete tests that he felt demonstrated the limitations of the GPT approach. GPT-3 smashed t
by andyljones 6y ago
Gary Marcus - the author of this - has previously offered several concrete tests that he felt demonstrated the limitations of the GPT approach.
GPT-3 smashed them.
https://www.gwern.net/GPT-3#marcus-2020 https://www.gwern.net/GPT-3#marcus-2020
- deleted 6y ago[deleted]
- lacker 6y agoFrom that link: Q. If a water bottle breaks and all the water comes out, how much water is left in the bottle, roughly? A. … Roughly half. … If the bottle is full, there is no water left in the bottle. I wouldn’t describe this as GPT-3 “smashing” the questions. It’s still clearly subhuman. This sort of question, logical real-world reasoning embedded in a descriptive sentence, is still hard for it. It’s definitely improving on GPT-2 though.
- abiro 6y agoOpenAI would naturally optimize for the tests published by Marcus as a critique of GPT-2, yet GPT-3 still fails physical reasoning spectacularly (the one test needing casual reasoning the most). There are two broader points here: 1. The lack of independently verifiable evaluation metrics for these type of models should make everyone very skeptical. (Who can afford to retrain GPT-3 from scratch?) 2. I find it difficult to believe that smart people still insist that a model incapable of representing causal relationships can produce intelligent answers.
- SpicyLemonZest 6y ago(1) I certainly agree with. But Marcus doesn't claim skepticism about GPT-3s intelligence; he claims that his evaluation metrics definitively show it doesn't understand the text it outputs or know anything about the world. (2) is, I think, a misunderstanding. People who believe GPT-3 is producing intelligent answers generally believe it can represent causal relationships.
- moyix 6y ago> OpenAI would naturally optimize for the tests published by Marcus as a critique of GPT-2 It would be difficult for them to do so since Marcus's GPT2 critique came out after they collected the dataset for GPT3. Marcus's article: Jan 2020 GPT-3 dataset: "Table 2.2 shows the final mixture of datasets that we used in training. The CommonCrawl data was downloaded from 41 shards of monthly CommonCrawl covering 2016 to 2019"
- Barrin92 6y ago>GPT-3 smashed them. which isn't surprising because virtually all of the questions are so simple they could literally appear in the training data that GPT-3 was trained on. I'm a little tired of proving how "intelligent" GPT is by asking these superficial questions. the MIT article gives much better examples that actually require physical, biological or higher-level reasoning and it produces complete nonsense as one would expect.
- Veedrac 6y agoThe article is meaninglessly cherry-picked, showing six bad answers out of 157, except those 157 examples were themselves cherry-picked to be bad out of a larger set. As usual, Gary Marcus is absurdly biased. For example, out of the larger 157 cherry-picked examples, there is this. > You poured yourself a glass of cranberry juice, but then absentmindedly, you poured about a teaspoon of grape juice into it. It looks OK. You try sniffing it, but you have a bad cold, so you can’t smell anything. You are very thirsty. So you drink it. It tastes a little funny, but you don’t really notice because you are concentrating on how good it feels to drink something. The only thing that makes you stop is the look on your brother’s face when he catches you. They then consider this a failure because, I quote, there is no reason for your brother to look concerned. This is patently ridiculous. It indicates that Gary has no idea what a language model even is. GPT-3 is not a Q&A model. It is not given a distinction between its prompt and its previous continuation. The only thing GPT-3 does is look for likely continuations. If you want GPT-3 to avoid story continuations, don't give it a story to continue! Or at least tell it what you're grading it on! But no, as usual, to Gary, all the times we show GPT-3 making sophisticated physical and biological deductions are fake, spurious, or meaningless. [1], [2], [3], [4]; none of that is truly evidence. But an incredibly cherry-picked, unfairly marked exam where you never told the examinee what you were testing them on, and you used high-temperature sampling without best-of, so only getting half right doesn't even indicate anything anyway (and of course, let's also pretend there are as many ways to be wrong as to be right, such that we can pretend each is equal evidence)—now that's enough evidence to write a disparaging article about how GPT-3 knows nothing. [1] https://twitter.com/danielbigham/status/1295864369713209351 https://twitter.com/danielbigham/status/1295864369713209351 [2] https://www.lesswrong.com/posts/L5JSMZQvkBAx9MD5A/to-what-extent-is-gpt-3-capable-of-reasoning https://www.lesswrong.com/posts/L5JSMZQvkBAx9MD5A/to-what-ex... [3] https://twitter.com/QasimMunye/status/1278750809094750211 https://twitter.com/QasimMunye/status/1278750809094750211 [4] https://news.ycombinator.com/item?id=23990902 https://news.ycombinator.com/item?id=23990902
- not2b 6y agoNo, those concrete tests are mostly issues that researchers have been talking about for years, meaning that many of them appear on the Internet somewhere. Increasing the volume of training data to hundreds of gigabytes likely meant that the exact questions and answers appeared in the training data. So GPT-3 didn't "smash them", it cut and pasted the answer from its training.