4 ms·
The idea that GPT-4 is 1 trillion parameters has been refuted by Sam Altman himself on the Lex Fridman podcast (THIS IS WRONG, SEE CORRECTION BELOW). These day
by techbruv 3y ago
The idea that GPT-4 is 1 trillion parameters has been refuted by Sam Altman himself on the Lex Fridman podcast (THIS IS WRONG, SEE CORRECTION BELOW).
These days, the largest models that have been trained optimally (in terms of model size w.r.t. tokens) typically hover around 50B (likely PaLM 2-L size and LLaMa is maxed at 70B). We simply do not have enough pre-training data to optimally train a 1T parameter model. For GPT-4 to be 1 trillion parameters, OpenAI would have needed to:
1) somehow magically unlocked 20x the amount of data (1T tokens -> 20T tokens)
2) somehow engineered an incredibly fast inference engine for a 1T GPT model that significantly better than anything anyone else has built
3) is somehow is able to eat the cost of hosting 1T parameter models
The probability that all the above 3 have happened seem incredibly low.
CORRECTION: The refutation for the size of GPT-4 on the lex fridman podcast was that GPT-4 was 100T parameters (and not directly, they were just joking about it), not 1T, however, the above 3 points still stand.
- sebzim4500 3y ago>The idea that GPT-4 is 1 trillion parameters has been refuted by Sam Altman himself on the Lex Fridman podcast. No it hasn't, Sam just laughed because Lex brought up the twitter memes.
- ftxbro 3y agonot sure why you're getting so downvoted lol
- tempusalaria 3y ago1) common crawl is >100TB so obviously contains more than 20trn tokens + Ilya has said many times in interviews that there is still way more data for training usage >10x 2) GPT-4 is way slower so this point is irrelevant 3) OpenAI have a 10000 A100 training farm that they are expanding to 2500. They are spending >$1mln on compute per day. They have just raised $10bln. They can afford to pay for inference
- CaptainNegative 3y ago> OpenAI have a 10000 A100 training farm that they are expanding to 2500. Does the first number have an extra zero or is the second number missing one?
- tempusalaria 3y agoSecond number is missing a zero sorry. Should be 10000 and 25000
- deleted 3y ago[deleted]
- nabakin 3y agoGPT-2 training cost 10s of thousands GPT-3 training cost millions GPT-4 training cost over a hundred million [1] GPT-4 inferencing is slower than GPT-3 or GPT-3.5 OpenAI has billions of dollars in funding OpenAI has the backing of Microsoft and their entire Azure infra at cost There is no way GPT-4 is the same size as GPT-3. Is it 1T parameters? I don't know. No one knows. But I think it is clear GPT-4 is significantly larger than GPT-3. For fun, if we plot the number of parameters vs training cost we can see a clear trend and I imagine, very roughly predict the amount of parameters GPT-4 has https://i.imgur.com/rejigr5.png https://i.imgur.com/rejigr5.png https://www.desmos.com/calculator/lqwsmmnngc https://www.desmos.com/calculator/lqwsmmnngc [1] > At the MIT event, Altman was asked if training GPT-4 cost $100 million; he replied, “It’s more than that.” http://web.archive.org/web/20230417152518/https://www.wired.com/story/openai-ceo-sam-altman-the-age-of-giant-ai-models-is-already-over/ http://web.archive.org/web/20230417152518/https://www.wired....
- cubefox 3y ago> There is no way GPT-4 is the same size as GPT-3. Is it 1T parameters? I don't know. No one knows. But I think it is clear GPT-4 is significantly larger than GPT-3. That's a fallacy. GPT-3 wasn't trained compute optimally. It had too many parameters. A compute optimal model with 175 billion parameters would require much more training compute. In fact, the Chinchilla scaling law allows you to calculate this value precisely. We could also calculate how much training compute a Chinchilla optimal 1 trillion parameter model would need. We would just need someone who does the math.
- nabakin 3y agoWhy does it matter in this case if GPT-3 was trained compute optimally or not? Are you saying that the over $100 million training cost is amount of training necessary to make a 175B parameter model compute optimal? And if they are the name number of parameters, why is there a greater latency with GPT-4?