4 ms·
Read though most of the paper and here's what GPT-3 is: If you wanted to generate poems with GPT-2, you'd need to have a lot of poems to fine-tune GPT-2 to get
by sdan 6y ago
Read though most of the paper and here's what GPT-3 is:
If you wanted to generate poems with GPT-2, you'd need to have a lot of poems to fine-tune GPT-2 to get reasonable results.
With GPT-3, you use few-shot learning instead (without the need to do gradient updates with each example)
The paper is long and filled with how it stacks with models like Grover and T5 and it does well... given that this is a 175 B param model (relative to Grover/T5's 1.5/11B param models). This shows that even with these huge models, smaller models can outperform them in certain instances with lesser param models.
Also I think they did a good job with explaning the ethics and morals around what models like these mean / what biases this has.
- ericlewis 6y agoWould you have any easy to explain insight in to how these perform better than larger models? I’ve always wanted to understand that as a technically adept and somewhat familiar (briefly) person who has explored what such models can do.
- sdan 6y agoThrow more computers and do some model architecture changes /s (although sometimes it's true)
- scottlocklin 6y ago>how these perform better than larger models they probably don't particularly; their inventors seem to excel in their PR budget rather than their verifiable innovations
- deleted 6y ago[deleted]
- Analog24 6y agoThe key insight in this paper is that the new (larger) model was not "fine-tuned" on the downstream NLP tasks. In other words, after it's trained on unsupervised (you could call it self-supervised in this case) data to do simple things like predict the next word (hence why it doesn't take any real supervision) it can then be used to do very specific tasks like answer questions or translating text without further supervision. Previous large-scale language models like BERT and GPT-2 had took a similar approach but in order to actually perform the more complicate down stream tasks they had to be fine-tuned. So they were trained with specific QA or translation date in order to understand and do well on those tasks. GPT-3 doesn't do any fine-tuning, it is able to take it's very general initial learning and perform very well on specific tasks that it was never trained on. This is why it doesn't perform as well as the "smaller" models on those tasks. But that is besides the point, if GPT-3 was fine-tuned on those tasks I'm sure it would achieve the latest SOTA results in many (all?) of them. The exciting part is how it was able to generalize the knowledge learned during "pre-training" to much more specific tasks. tl;dr the smaller models were trained on the specific tasks that they were evaluated on. The large model (GPT-3) was not trained on those specific tasks and still does almost as well.
- ericlewis 6y agovery cool, thanks for explaining!
- drdeca 6y agoHave you tried making GPT2 do zero-shot poetry writing? It's not great at it, but it is good enough at it to get something interesting enough if you try a number of times. Go to talktotransformer.com/ and give it the prompt "Here is a poem I wrote:" or "Here is my favorite poem:" . I'm sure GPT3 would produce much better and more consistent results, but GPT2 will produce something that looks generally like a poem frequently enough, and sometimes it will even be relatively coherent? Here is one that it produced for me today: > You say, "Don't lose your pride." > Here is my rejoinder: > Well, maybe it is the pride of a diseased soul. > You are a wanderer, you know not whence, > O thief, you fool, you rhinoceros > Caught in the jaws of a viper. > You may lament your affliction > For the world will laugh at your tears. > Pray to a demi-god > Hail him and say, > "Ah, Sir, give me thy pity! > O thou who maintainest as if thou wert a king! > Here is thy axe, I say; let us Is it great? No. But it has some level of coherence. Here is another: Let me tell you the reason I love poetry. // All the things of the world I have described, // If you ask me why I like poetry, // It would seem quite simple to me. // When I'm working at the computer in the evening // I'll get out my books of poems and I'll turn them over, // Like blades of grass under the hot sun, // That write with such fineness the kind of green I like. // But if I'm a bit more tired in the morning, // I'll fill a little stack of yellow pages with poems, // That let the air and the dry light of morning run wild. // You know, the// This one even rhymes a bit!: the sword is to slay // The axe is to smite // The stick is to break // The tooth is to bite // All these are for our earthly security, // All have their uses, // Those which can be used // Must be employed. // The sword is the instrument of strife // The axe is the weapon of war // The stick is the weapon of domestic strife // The tooth is the instrument of war // All these be in our hands. // At the time of our death // They will be in our hands, // And then we will weep, // Though now we sleep. // —The Remaining Three Questions (sorry, idk how to format these to make them look right. The leading "> " and the "//" insertions are me trying to format them to make the line separators clear.)
- stared 6y agoNot for poems, but for AI-generated cities (inspired by "Invisible Cities" by Italo Calvino), e.g.: PROMPT: Cities & Lights When you enter the city of Singapore during the night, you see lights: colorful and ubiquitous. Lights on every building, on every fountain, and in every park. GENERATED: Lights shining in a city in which the majority of people are now using mobile phones. Singapore has a bright future as a technology hub, and it 's not too late to make it happen. ... On the other occasions I seeded with ~two sentences of "Invisible cities", and it worked like a charm, no fine-tuning.