7 ms·
GPT3's strength is on language generation, so using *GLUE for evaluating it (where encoder type models are just better) and claiming to have 99.9% less paramete
by lzzzfelipe 6y ago
GPT3's strength is on language generation, so using *GLUE for evaluating it (where encoder type models are just better) and claiming to have 99.9% less parameters is sensationalism.
- visarga 6y agoYes, it can do simple yes/no questions but can it write dad jokes like GPT-3?
- dvfjsdhgfv 6y agoFor me another question is crucial: is it open (i.e. is there an open source reference implementation) or closed like the so-called "OpenAI" products/services?
- mkolodny 6y agoYou can find the code for the paper here: https://github.com/timoschick/pet https://github.com/timoschick/pet
- emteycz 6y agoCan you please explain to me (I don't understand it) why do you think GPT-3 is closed? Yes, they won't share the trained model, but they're sharing the research here[0][1] so you can reproduce easily, aren't they? As I understand it now, it's very fair - training the model is a separate thing from doing (and sharing) the research, is very costly, and would not happen if they were forced to open that too - I also don't understand why should they be. [0] https://arxiv.org/abs/2005.14165 https://arxiv.org/abs/2005.14165 [1] https://github.com/openai/gpt-3 https://github.com/openai/gpt-3
- dj_mc_merlin 6y ago> so you can reproduce easily That's like saying you can look at the Eiffel tower and it's schematics so what's so hard about getting a spare million dollars and building it.
- emteycz 6y agoWell, we others can buy a ticket and go see it. Same with GPT-3.
- fxtentacle 6y agoYes they tell you how to do it. But training a model of this size requires you to use thousands of GPUs, or wait forever. That will sum up to millions of $$$ in rental and electricity costs.
- emteycz 6y agoThat seems like the reason why won't they share the trained model. Somebody needs to pay for that, how would they fund it if they just shared it? As far as my thinking goes, they are open more than enough. Thank you for your input!
- fxtentacle 6y agoThey claim to be a well funded nonprofit.
- emteycz 6y agoYeah, and that nonprofit probably wouldn't really be doing its mission it got funding for if they were handing out computing, instead of research
- fxtentacle 6y agoWell, they are withholding the information needed to verify their research. That is generally frowned upon in science. Plus, they are called OpenAI but producing a closed source product...
- emteycz 6y agoIt's not a product, it all boils down to raw computing work. Their mission is to do open research, not provide computing time that they paid for for free. Similarly, Red Hat does not have to provide free cloud computing because they also contribute to Linux and use it.
- dvfjsdhgfv 6y ago
- luma 6y agoI submitted a request for access to the GPT-3 API a couple months ago and still haven't been approved. I don't find that to be very open at all.
- emteycz 6y agoThe API does not have to be open, the research does. This is like saying cloud services should be free and open for all because Linux is open.
- emteycz 6y agoTo expand on that point, they're basically selling pre-executed computing to you. Resolving the query itself is nothing.
- dvfjsdhgfv 6y agoFor quite some time maddog has been promoting Subutai, a kind of P2P cloud (https://subutai.io/ https://subutai.io/). Voluntary computing is an old concept, even BOINC is almost 20 years old. Peertube already reached the stage where it's actually usable and people can watch the movies smoothly. So it's not unimaginable that people who care organize somehow creating a platform where you could use a GPT-3 platform by contributing your GPU time in exchange.
- dynamite-ready 6y agoThis is a very article on the subject of GPT3's 'market position' - https://bdtechtalks.com/2020/09/21/gpt-3-economy-business-model/ https://bdtechtalks.com/2020/09/21/gpt-3-economy-business-mo...
- gvhst 6y agoThat's kind of like saying that this super cheap SUV from 2003 is a better off-roader than the latest Ferrari. Like, true statement, but vacuous none the less.
- tinmandespot 6y agoOh I love me a great analogy!!
- leereeves 6y agoPowerful analogy, but analogies are dangerous. They can obscure what's really happening. In this case, by analogy, Ferrari made the comparison to the super cheap SUV from 2003. That is, OpenAI compared GPT3 to BERT on the SuperGLUE benchmark, in the paper announcing GPT3 [1]. They did so to demonstrate GPT3's ability to learn a new task given only a few examples of the task ("few-shot learning"). The limited amount of task-specific training data was a signficant handicap that GPT3 was able to overcome, like a Ferrari towing a two ton trailer outperforming an old SUV towing nothing. What this paper claims is that encoder type models can also achieve few-shot learning. The headline should be "AI training method achieves few-shot learning with 99.9% fewer parameters than GPT3." That's the innovation here, not outperforming GPT3 on a benchmark that GPT3 isn't particularly good at. 1: https://arxiv.org/pdf/2005.14165.pdf https://arxiv.org/pdf/2005.14165.pdf
- Cybiote 6y agoUnfortunately, your correction is still misleading because it fails to capture what really sets GPT-3 apart. Ironically, GPT-3's limitation also highlights its strength. As it is not capable of learning (few shot or otherwise) in the strict sense of permanently changing its parameters based on examples, all its demonstrated capabilities are completely at inference time. It somehow configures itself at inference time so that state machines which produce plausible continuations of whatever pattern it was fed, are most probably generated. This means that whenever it succeeds, it is much more flexible in how it produces its responses. It generalizes on and continues those implicit patterns in the provided input. This paper however, is not replicating that flexibility. Their proposed model is much closer to expectation maximization than it is an instance of what GPT3 does. The novelty of their work, what makes it genuinely useful, is they provide a practical and fairly general way to leverage pre-trained language models to produce classifiers for specific tasks using a very small amount of labeled data. Requiring less effort compared to what would go into fine-tuning. This approach to distillation is an instance of https://en.wikipedia.org/wiki/Semi-supervised_learning https://en.wikipedia.org/wiki/Semi-supervised_learning. Compared to GPT-3, this approach remains at a severe disadvantage when amount of effort and time required to gather data and train something useful is accounted for. On the other hand, if you can fit your problem into the format proposed by the paper, you will likely have more control on the final model's behavior using a small amount of labeled examples (for your specific task), at a significantly lower cost of computation at inference time. Focusing on parameters however, is not even wrong.
- nl 6y agoThis isn't strictly true. This is comparing to "GPT3 as a few-shot learner"[1] as opposed to the fine tuned models that everyone else use. Few-shot GPT3 outperforms a BERT-based baseline. [1] https://github.com/openai/gpt-3 https://github.com/openai/gpt-3
- empiko 6y agoYeah, but they actually do fine-tune their model and compare it to in-batch GPT3 "learning". I called them out on it on Reddit and they claim that they use the same amount of data. I.e. GPT use X examples in the sample and they use the same amount of X samples to fine-tune. However, I am still not convinced that it is a fair comparison.
- fnord123 6y agoWelcome to benchmarklandia, where you find the top algorithm in an area and write a stripped down version that outcompetes it on a different set of data and get good results. See: hadoop->spark->spark with infiniband->spark with nvme->timely dataflow on a laptop; csv parsers that don't support scientific notation floats, etc. I'm sure others can give moren interesting examples.
- albertwang 6y agoOut of (self-interested) curiosity, what are you referring to with the "nvme->timely dataflow on a laptop" reference? The closest benchmark I could find was the "FASTER State Management for Timely Dataflow" paper from ETH Zurich, but that wasn't run on a laptop.
- fnord123 6y agohttp://www.frankmcsherry.org/assets/COST.pdf http://www.frankmcsherry.org/assets/COST.pdf
- frankmcsherry 6y agoIt's a bit different, right, because that paper presents better algorithms on the same datasets as the work it cites.
- fnord123 6y agoYes it's different. Another part of Benchmarklandia is where researchers make custom algorithms for the same dataset. Benchmarks targeting MNIST is a great example. :)
- Veedrac 6y agoI had this discussion with Timo Schick, the first author, on Reddit. My final comment is copied below, with the relevant context. Note that GPT-3's approach to *GLUE involved no training on the task, just a good choice of prompt, whereas PET and iPET also use fine-tuning. Also, because distillation takes large amounts of training data, they use ensembles in the true few-shot regime, so their parameter efficiency is significantly worse than they advertise. https://www.reddit.com/r/slatestarcodex/comments/itrcac/small_language_models_are_also_fewshot_learners/g5p0dl5/ https://www.reddit.com/r/slatestarcodex/comments/itrcac/smal... --- Timo Schick: Finally, I do not really agree with your last two paragraphs, especially "One is about semi-supervised learning, that says by exploiting task-specific architectures you can do fairly well with low amounts of labelled data.": If you leave out the final distillation step (which is not required for good performance), we use the exact same architecture for all tasks. In what sense is this more task-specific than GPT-3? I would not consider "exploiting task-specific architectures" to be a (fundamental) part of the paper. My reply: So what I mean here is that masked training and bidirectional transformer models like BERT have always been designed as a way to get good scores in analysis tasks like Q&A, even if they are pretrained on general text, whereas unidirectional generative transformer models are now basically only relevant for generative tasks. You can say, well, both architectures can do both tasks, so is it really task specific?, but ultimately, yes, we've selected ALBERT because it's better for Q&A tasks, and we've selected unidirectional transformers in other things because they're better for generative tasks. So I guess the problem I have is with the merits of your thesis, “Can we achieve similar few-shot performance to GPT-3 without requiring billions of parameters?” OpenAI didn't present few-shot learning as if it were an optimal method; their headline achievement was not “here's the best way to...” but “I bet you never expected that this could...”. And so while it's definitely true that a BERT-derived model will outperform a GPT-derived model even at lower parameter counts on these sort of tasks, nothing new or interesting is being said by it. Everyone already knows that a bidirectional GPT-3 would be better at Q&A, and so that's what a smaller bidirectional model should be competing against. GPT-3 is only interesting in this context because it's not the optimal model (or training routine). So while it's also true that if your aim is SOTA in few-shot learning then you should definitely use a bidirectional transformer with all the new tricks, if your goal is to understand PET in a context that includes GPT-3, doing so merely makes it harder to see what's going on.