4 ms·
Kind of. This tweet better explains what they did: https://twitter.com/abhi_venigalla/status/1673813863186452480?s=20 https://twitter.com/abhi_venigalla/status
by april7 3y ago
Kind of. This tweet better explains what they did:
https://twitter.com/abhi_venigalla/status/1673813863186452480?s=20 https://twitter.com/abhi_venigalla/status/167381386318645248...
- smaddox 3y agoSo it would be more like 46 hours to train GPT-3 from scratch. (300BT / 1.2BT) * 11 min * (1 hr / 60 min) = 45.8 hr Still pretty incredible. That's an 18.8x speedup over 36 days.
- ChatGTP 3y ago[flagged]
- deleted 3y ago[deleted]
- lumost 3y agoThis points to an interesting future for foundation models. This is an 18x cost reduction in only 2 years. Either foundation models are going to get much bigger, or variations will become common.
- rerx 3y agoV100 GPUs are from 2017, so it's more than two years. A100 already appeared there years ago, btw. An eight GPU DGX-1 server cost ~149k$ back then (googled news postings). A current gen DGX H100 is 520k$ with 5 years of support. Of course it holds 5x the memory, plus GPUs and interconnect are much faster. But when comparing costs, take price hikes into account.
- jsjohnst 3y agoAn important thing to also keep in mind is how much inflation changed prices over the duration. $520k in 2023 dollars is around $420k in 2017 dollars. Sure, still almost 3x more expensive, but that’s better than being 0.7x higher.
- raverbashing 3y agoVariations of specializations I guess For writing code you don't care about feeding world history to your model. So a smaller model might be better at a specialized task Sure, having a big multi-modal-model is great, but by having specialized models you can spread tasks better
- mlboss 3y agoBut I am sure prompt understanding improves with more text data. Same with reasoning ability.