4 ms·
Disclaimer: I know absolutely nothing about machine learning. Isn't GPT-3 the architecture? Are they doing something different or why would it not scale?
by ImprobableTruth 6y ago
Disclaimer: I know absolutely nothing about machine learning.
Isn't GPT-3 the architecture? Are they doing something different or why would it not scale?
- m00x 6y agoIt's the model, not the architecture, but you could say the model contains the architecture.
- nmfisher 6y agoGPT-3 is the name for the architecture, but there are a few different versions/sizes. The OpenAI version that impressed us all was ~170B parameters, this is far smaller. To go from 2.7B to 170B parameters will need more than just a few config tweaks. There's a whole bunch of hacks and tricks needed to coax a model to train at that scale, the Eleuther version is almost guaranteed to fail out-of-the-box.
- minimaxir 6y agoIt's worth noting that the GPT-3 paper did train models with more sane sizes (e.g. 1.5B) as a point of comparison. I am surprised/annoyed they never released them though.
- stellaathena 6y agoIt's because OpenAI sells them for profit. The "Ada" model is the same size as the larger of these two EleutherAI models.
- minimaxir 6y agoHuh, I was wondering what the size of the non-davinci models were; guess that make sense. It's still telling that a "small" GPT-3 model can risk cannibalizing a larger model.
- stellaathena 6y agoAda is 2.7B, Babbage is 6.7B, Curie is 13.0B, and DaVinci is 175B. The new one they announced last month is in the 20-50B range I think, not totally sure though.
- sendtown_expwy 6y agoI would guess that an average FAANG ML engineer could code up and successfully execute a forward/backward pass on a GPT-1 or GPT-2 model with a day of effort or less. (GPT-3 a little harder, but not significantly). But is that model actually going to perform well? Most likely no. Model performance varies significantly due to subtle details in data processing implementations, seemingly insignificant details in code, and even from different numerical methods of calculating the same semantics. If you don't believe me, consider that many ML researchers track their commits (or exact code versions) extremely carefully, because oftentimes they will make some change (or changes) they think are inconsequential and later find that actually, their model broke. If they made too many changes, whoops, guess you have to binary search over the diff to see what happened since your last "good run". If the people who spent months (if not years) tuning a model can't tell whether it will work from the code, how could anyone else? Most ML researchers will not bother with most code that doesn't give proof of results (in terms of a model that can actually be evaluated) because it is just so unlikely that it will actually work well. Now, it might "work" in the sense that it converges and does something when you prompt it with examples. But will this GPT-3 reimplementation actually outperform say, the 10x smaller T5 checkpoint that was released by Google, or the other smaller language models others have released? If it doesn't, it's hard to argue that its very useful at all. I think that's the spirit of why the original commenter said what they did, but I still do applaud the efforts of this team (and hope that their implementation is, in fact, highly performant!)
- nl 6y agoAlmost all the challenges with GPT-sized models are engineering and training challenges, not architectural. How do you train a model too big to fit in a single GPU? It's doable, but not simple. How do you update weights across your cluster? etc etc