3 ms·
Is the word "GPT-3" now being used like a a generic term, simply referring to a large-enough language model? From the article, Meta's OPT model seems to be inte
by needle0 4y ago
Is the word "GPT-3" now being used like a a generic term, simply referring to a large-enough language model? From the article, Meta's OPT model seems to be intentionally designed to match the capabilities of OpenAI's original GPT-3, but doesn't say anywhere that it shares any code lineage.
- nodja 4y agoThat's because the code is unimportant. GPT-2 code is very simple, I think the official release by openai is only about a couple hundred lines of code. The challenge of GPT-3 is the scale, the GPT-3 paper basically says "This is GPT-2, but we made changes in the model so we can run it on dozens of GPUs". The changes mostly don't matter, because if you had a big enough GPU (~1.5TB VRAM) you could just up the hyperparameters of GPT-2 and you'd end up with the same results. So what's novel about GPT-3 that warrants a new name? The discovery here is that after the models reach a certain size, it's able to do many tasks without any training. You can literally ask it to translate from one language to another, at it'll do a decent job at it, if you give it a few examples it'll work even better. Now that doesn't mean that GPT-3 is the final model, it's still not good enough for many tasks, for example copilot is based on GPT-3 but it was specifically fine-tuned for the task of auto-completing code. So yes, if you can figure out how to scale a GPT-2 model to 175B parameters, you have a GPT-3 clone.