11 ms·
GPT-3: A Disappointing Paper?
- The_rationalist 6y agoMeanwhile, a revolutionary paper that brought for the first successful time a new paradigm to NLP (latent variational autoencoders) and that destroy GPT 3 on text perplexity on the Pen treebank (4.6 vs 20) and with order of magnitudes less parameters is talked about nowhere on the web... https://arxiv.org/abs/2003.02645v2 https://arxiv.org/abs/2003.02645v2
- mchusma 6y agoI've seen you post this before and you seem enthusiastic about it. Do you have a good summary of why you feel this paper so important?
- The_rationalist 6y agowhy you feel this paper so important? the best summary are the benchmarks: first place at question answering (on yahoo task) First place on language modeling for pen treebank (by FAR) So it is a totally new model that will probably keep evolving and being applied to more and more kind of NLP tasks. And it seems that it can have the first place on most NLP tasks, its empirically the breakthrough of the year. It achieve this while having an extremely small number of parameters which shows: The model is smarter The model has room for more parameter hence even more accuracy! Finally, theoretically it is a breakthrough as it is a port of a computer vision technology (variational autoencoders) to the NLP world. Actually it might be the successor to the Transformer paradigm. I wonder what pile of incremental improvement will researchers will be able to bring to it like they have on the Transformer paradigm (spanBert, XLnet, etc)
- phreeza 6y agoNot all models scale up well with more parameters. VAEs are not really exclusively a vision thing, they have been used in a variety of settings. Using VAEs for NLP is also nothing new, an early example is Bowman et al, 2015. https://arxiv.org/abs/1511.06349 https://arxiv.org/abs/1511.06349
- The_rationalist 6y agoWhat's new is that they are the first to eliminate the bottleneck that prevented VAES to beat transformers
- phreeza 6y agoWhat kind of bottleneck are you talking about? VAEs as such are really in a different category from transformers. They are primarily a tool to get to more structured latent spaces, which is not something transformers are good at in the first place.
- milquetoastaf 6y agoAny models or code available?
- The_rationalist 6y agohttps://github.com/seraphlabs-ca/SentenceMIM-demo https://github.com/seraphlabs-ca/SentenceMIM-demo
- laretluval 6y agoIt’s only been a few months. If it really does robustly get 4.6 perplexity on PTB (less than I ever thought was possible) then it will receive its due recognition, at the very least via the Hutter prize.
- psb217 6y agoFYI, the MELBO bound in that paper is invalid. Their perplexity numbers using the MELBO bound are also invalid.
- jaredtn 6y agoCan you explain further?
- psb217 6y agoI've added a brief explanation in a reply to a sibling comment.
- The_rationalist 6y agoWhere is the error? How much would that change their score of 4.6?
- psb217 6y agoThe bound is completely invalid, as are the NLL/PPL numbers they report with the MELBO. Look at the equation. If they optimized it directly, it would be trivially driven to 0 by the identity function if we used a latent space equivalent to the input space. The MELBO just adds noiseless autoencoder reconstruction error to a constant offset equal to log of the test set size. This can be driven to zero by evaluating an average bound over test sets of size 1. The mathematical/conceptual error is that they are assuming each test point is added to the "post-hoc aggregated" prior when they evaluate the bound. This is analogous to including a test point in the training set. Another version of this error would be adding a kernel centered on each test point to a kernel density estimator prior to evaluating test set NLL. In this case, obviously the best kernel has variance 0 and assigns arbitrarily high likelihood to the test data.
- The_rationalist 6y agoInteresting, thx!
- lonelappde 6y agoHave they deployed a demo? It's hard to talk about something that no one can see.
- The_rationalist 6y agoThey have: https://github.com/seraphlabs-ca/SentenceMIM-demo https://github.com/seraphlabs-ca/SentenceMIM-demo It's far easier to reproduce than GPT-3 because you won't need a GPU farm, only a powerful one
- canjobear 6y agoThe perplexity numbers are for different tasks. MIM is encoding a sentence into a latent variable and then reconstructing it, and achieves PTB perplexity 4.6. GPT-2 is generating the sentence from scratch, which will on average have higher perplexity numbers. I agree that PTB perplexity 4.6 on autoregressive language modeling would be a huge result.
- p1esk 6y agoI think using AE for text generation is a good idea, and is pretty old one (I tried it myself back in 2015 without particularly good results), but I wouldn't call it a breakthrough. To me a breakthrough/new paradigm would be something like this: https://arxiv.org/abs/1906.05317 https://arxiv.org/abs/1906.05317
- The_rationalist 6y agoThe breakthrough is not that they are the first to use AE for text. It is that they are the first to eliminate the bottleneck that prevented AEs to beat transformers. You linked paper seems interesting on paper but does it bring any new SOTA?
- p1esk 6y agoAEs to beat transformers Can you point me to some example of generated text this model produced? Something similar in quality to that unicorn story from GPT-2? but does it bring any new SOTA? Looking at their tables, seems so. The code is open source, and there's demo at https://mosaickg.apps.allenai.org/ https://mosaickg.apps.allenai.org/
- strin 6y ago> “GPT-3″ is just a bigger GPT-2. In other words, it’s a straightforward generalization of the “just make the transformers bigger” approach Yes it’s true. But there is a difference between what’s interesting and what works. deep learning (RNNs, transformers, etc.) is usually old ideas applied at large scale with slight modifications. Proving a model works well at large scale (175B parameters) is a great contribution and measures our progress towards AI.
- cs702 6y agoAll valid points, but I disagree with the conclusion, for several reasons: * First of all, the GPT-3 authors successfully trained a model with 175 billion parameters. I mean, 175 billion. The previous largest model in the literature, Google’s T5, had "only" 11 billion. Models with trillions of weights are suddenly looking... achievable. That's a significant experimental accomplishment. * Second, the model achieves competitive results on many NLP tasks and benchmarks without finetuning, using only a context window of text for instructions and input. There is only unsupervised (i.e., autoregressive) pretraining. AFAIK, this is the first paper that has reported a model doing this. It's a significant experimental accomplishment that points to a future in which general-purpose NLP models could be used for novel tasks without requiring additional training from the get-go. * Finally, the model’s text generation fools human beings without having to cherry-pick examples. AFAIK, this is the first paper that has reported a model doing this. It's another significant experimental accomplishment. More generally, I find that some AI researchers and practitioners with strong theoretical backgrounds tend to dismiss this kind of paper as "merely" engineering. I think this tendency is misguided. We must build giant machines and gather experimental evidence from them -- akin to physicists who build giant high-energy particle colliders to gather experimental evidence from them. I'm reminded of Rich Sutton's essay, "The Bitter Lesson:" http://www.incompleteideas.net/IncIdeas/BitterLesson.html http://www.incompleteideas.net/IncIdeas/BitterLesson.html
- 100721 6y agoI mostly agree with you. However, > It's a significant experimental accomplishment that points to a future in which general-purpose NLP models could be used for novel tasks without requiring additional training from the get-go. This premise is still purely science fiction. This model does not touch on either novel tasks nor being free from pretraining (unless I misunderstand). But overall, I think you’re right: it’s significant for a number of reasons.
- cs702 6y agoGPT-3 was pretrained on five datasets (Common Crawl, WebText2, Books1, Books2, and Wikipedia; see table 2.2), and then used on previously unseen tasks (Q&A, translation, cloze, etc.) without finetuning, i.e., weights were not updated after the original (autoregressive) pretraining. This promises a possible future in which general-purpose models are pretrained once, and deployed to production for multiple tasks.
- Voloskaya 6y ago> Transformers are extremely interesting. And this is about the least interesting transformer paper one can imagine in 2020. Because it's not a transformer paper. This paper goal was to see how far can an increase in compute continue to deliver an increase in model performance. There is no better way to study this than to take a very well known architecture and keep it the same as possible, otherwise it becomes very hard to know what is due to the increase size of the model and what is due to the tweaks you make. So yes, it's a disappointing paper if you expect it to be on a different topic than what it is.
- deleted 6y ago[deleted]
- krzyk 6y agoWhat does GPT mean? I assume it is not about partition tables (GUID Parition Table), it has something to do with NLP, but besides that it is hard to find what does this acronym mean.
- jaredtn 6y agoGeneralized Pretraining. The original GPT paper doesn’t use the acronym, but got rebranded as GPT in retrospect once GPT-2 came out.
- Smaug123 6y agoSupposedly it stands for "Generative Pretrained Transformer", but nobody ever expands the acronym; it's a language model, originally released by OpenAI and announced at https://openai.com/blog/language-unsupervised/ https://openai.com/blog/language-unsupervised/ and https://openai.com/blog/better-language-models/ https://openai.com/blog/better-language-models/ .
- reddickulous 6y agoIt would be cool if there was a platform to crowd source compute resources to train stuff like this so that regular people (without 7 figure budgets) can have access to these models which are becoming increasingly out of reach to the general public.
- deleted 6y ago[deleted]
- mryab 6y agoHere is a recent paper (disclaimer: I am the first author) named "Learning@home" which proposes something along these lines. Basically, we develop a system that allows you to train a network with thousands of "experts" distributed across hundreds or more of consumer-grade PCs. You don't have to fit 700GB of parameters on a single machine and there is significantly less network delay as for synchronous model parallel training. The only thing you sacrifice is the guarantee that all the batches will be processed by all required experts. You can read it on ArXiv https://arxiv.org/abs/2002.04013v1 https://arxiv.org/abs/2002.04013v1 or browse the code here: https://github.com/learning-at-home/hivemind https://github.com/learning-at-home/hivemind. It's not ready for widespread use yet, but the core functionality is stable and you can see what features we are working on now.
- drcode 6y agoNewbie question: If/when models the size of GPT3 are released to the general public, will average people going to be able to run them on their PCs, as they can with GPT2? Or will that basically be impossible now without expensive specialty hardware?
- freeone3000 6y agoGPT-2 takes 500ms per word on our benchmarks on a Xeon 4114, compared to 15ms on a Titan RTX. So the answer is technically yes, practically no, but why would you?
- 6gvONxR4sf7o 6y agoThe big one is 175 billion parameters. With your hardware's usual 32 bit floats, that's a 700GB model. You won't be using the big one for a while.
- p1esk 6y agoThis one uses FP16, so you just need to have a server with >350GB of RAM. 512GB of DDR4 would set you back around two grand. A total cost of a server for this would probably be under $5k. Comparable to a good gaming rig.
- sillysaurusx 6y agoA TPU can allocate 300GB without OOMing on the TPU's CPU. That's tantalizingly close to 350GB. And 300GB + 8 cores * 8GB = 364GB. It'll take some work, but I think I can come up with something clever to dump samples on a TPUv2-8. i.e. the free one that comes with Colab. Realistically, I don't think OpenAI will release the model. Why would they? And I'm not sure they'd dare use "it might be dangerous" as an excuse.
- p1esk 6y agoHave you (or anyone) tried running GPT-2 inference in INT8 precision? Perhaps worth looking at one of these efforts: https://www.google.com/search?q=running+transformer+in+int8&oq=running+transformer+in+int8 https://www.google.com/search?q=running+transformer+in+int8&...
- ericjang 6y agoI could not disagree more with this post. To summarize what the author is unhappy with: 1) "It’s another big jump in the number, but the underlying architecture hasn’t changed much... it’s pretty annoying and misleading to call it “GPT-3.” GPT-2 was (arguably) a fundamental advance, because it demonstrated the power of way bigger transformers when people didn’t know about that power. Now everyone knows, so it’s the furthest thing from a fundamental advance." 2) "The “zero-shot” learning they demonstrated in the paper – stuff like “adding tl;dr after a text and treating GPT-2′s continuation thereafter as a ‘summary’” – were weird and goofy and not the way anyone would want to do these things in practice... They do better with one task example than zero (the GPT-2 paper used zero), but otherwise it’s a pretty flat line; evidently there is not too much progressive “learning as you go” here." 3) "Coercing it to do well on standard benchmarks was valuable (to me) only as a flamboyant, semi-comedic way of pointing this out, kind of like showing off one’s artistic talent by painting (but not painting especially well) with just one’s non-dominant hand." 4) "On Abstract reasoning..So, if we’re mostly seeing #1 here, this is not a good demo of few-shot learning the way the authors think it is." --------- My response: 1) The fact that we can get so much improvement out of something so "mundane" should be cause for celebration, rather than disappointment. It means that we have found general methods that scale well and a straightforward recipe for brute-forcing our way through solutions we haven't solved before. At this point it becomes not a question of possibility, but of engineering investment. Isn't that the dream of an AI researcher? To find something that works so well you can stop ``innovating'' on the math stuff? 2) Are we reading the same plot? I see an improvement after >16 shot. I believe the point of that setup is to illustrate the fact that any model trained to make sequential decisions can be regarded as "learning to learn", because the arbitrary computation in between sequential decisions can incorporate "adaptive feedback". It blurs the semantics between "task learning" and "instance learning" 3) This is a fair point actually, and perhaps now that models are doing better (no thanks to people who spurn big compute), we should propose better metrics to capture general language understanding. 4) It's certainly possible, but you come off as pretty confident for someone who hasn't tried running the model and trying to test its abilities. Who is the author, anyway? Are they capable of building systems like GPT-3?
- master_yoda_1 6y ago
- gambler 6y agoArticle> One of their experiments, “Learning and Using Novel Words,“ strikes me as more remarkable than most of the others and the paper’s lack of focus on it confuses me. This sort of "learning" is not necessarily real learning and it's not new for GPT-3. Even reduced GPT-2 willingly used made-up terms from the prompt in its results: https://medium.com/@VictorBanev/interrogating-gpt-2-345m-aaff8dcc516d https://medium.com/@VictorBanev/interrogating-gpt-2-345m-aaf... Search the article for 'Now I will feed it the same thing, but with a bunch of made-up terms.' It has some examples of how that stuff worked. I've already posted this in the original discussion of GPT-3 paper and I will post it again: statements about whether some system "learns new words" or "does math" require hypothesis formulation and testing. It astounds me that many people in ML community not only don't do these sort of things, but even actively oppose to the very idea of them being necessary. Recently there was a great live-stream from DarkHorse talking about this problem in science in general: https://www.youtube.com/watch?v=QvljruLDhxY https://www.youtube.com/watch?v=QvljruLDhxY They talk about "data-driven" science and the fundamental problems with that notion.
- bitL 6y agoI think the main disappointment is that we humans aren't that special when a brute-forced scalable transformer is getting into our ballpark. We have also recently seen how Open AI + MS were able to use a GPT-variation for automated text-description-to-python-code generation, and utilizing something like GPT-3 in that task might render many swengs obsolete fairly soon.
- victor9000 6y agoMy biggest problem with GPT3 is that it's not going to be accessible (practically speaking) to the general public. There's been a recent push to democratize this type of work with libraries like Huggingface transformers, but models this large will force the benefits of this work back into the ivory tower.
- jkhdigital 6y ago> it would represent a kind of non-linguistic general intelligence ability which would be remarkable to find in a language model As a relative outsider to this field, I don’t really see the stark line between natural language and general intelligence implied by this statement. Language is just abstractions encoded in symbols, and general intelligence is just the ability to construct and manipulate abstractions. Seems reasonable to think that these are two sides of the same coin. Put another way, natural language is the product of general intelligence.
- remexre 6y agoI think the line you'd see is that there exists some task where the language-based model suddenly lacks the ability to perform the task despite the fact that it "should." I'd conjecture that this might include something like describing where places are in relation to each other, and asking it to describe a route. (Not an NLP expert, but work with AI folks; this task chosen as an example because it seems like something you'd want a planner for rather than anything MLful.)