7 ms·
These could train GPT-3 in 46 hours. (Edited from "11 mins" per https://news.ycombinator.com/item?id=36500154 https://news.ycombinator.com/item?id=36500154 ) A
by ed 3y ago
These could train GPT-3 in 46 hours. (Edited from "11 mins" per https://news.ycombinator.com/item?id=36500154 https://news.ycombinator.com/item?id=36500154 )
An equivalent number of V100's (GPT-3's original GPU) would've taken about 36 days [0].
0 – https://www.reddit.com/r/GPT3/comments/p1xf10/comment/h8h3sl4/?utm_source=reddit&utm_medium=web2x&context=3 https://www.reddit.com/r/GPT3/comments/p1xf10/comment/h8h3sl...
- messe 3y ago3500 of them trained GPT-3 in 11 minutes. That number is still worth remarking.
- reaperman 3y agoIt’s remarkable that one of these could train GPT-3 in under a month. However, good to note that its the price is that of twenty 4090’s.
- ClassyJacket 3y agoCould it actually? Or do you need the memory size of thousands of them running in parallel?
- reaperman 3y agoYes I worded that extremely poorly.
- cavisne 3y agoOne interesting thing is if the model can’t fit into GPU memory (sharded across multiple chips) it would be much slower. So one chip would take more than a month. That’s why these clusters are so valuable, even with Nvidias margins they are still cheaper than using less compute for longer.
- visarga 3y agotitle says: "a massive GPT-3-based benchmark" that means they might pass a few batches through a GPT-3 sized random init network and time it
- layoric 3y agoGPT3 benchmark, as others have said this means it would be ~46 hours to train GPT, still impressive! Also, going by other comments here, ~3500 of these or 13-14 DGX GH200s at ~$10m each means we are talking about ~$135m worth of compute here. Still very impressive but holy hell that is a lot of money worth of compute hardware.
- zetazzed 3y agoYeah, it's wild - feels like we see "software is so slow now! nobody optimizes software anymore, they just run Python and burn cycles!" posts all the time, but man, when a company REALLY wants to optimize something -- it's a thing of beauty.
- april7 3y agoKind of. This tweet better explains what they did: https://twitter.com/abhi_venigalla/status/1673813863186452480?s=20 https://twitter.com/abhi_venigalla/status/167381386318645248...
- smaddox 3y agoSo it would be more like 46 hours to train GPT-3 from scratch. (300BT / 1.2BT) * 11 min * (1 hr / 60 min) = 45.8 hr Still pretty incredible. That's an 18.8x speedup over 36 days.
- ChatGTP 3y ago[flagged]
- deleted 3y ago[deleted]
- lumost 3y agoThis points to an interesting future for foundation models. This is an 18x cost reduction in only 2 years. Either foundation models are going to get much bigger, or variations will become common.
- rerx 3y agoV100 GPUs are from 2017, so it's more than two years. A100 already appeared there years ago, btw. An eight GPU DGX-1 server cost ~149k$ back then (googled news postings). A current gen DGX H100 is 520k$ with 5 years of support. Of course it holds 5x the memory, plus GPUs and interconnect are much faster. But when comparing costs, take price hikes into account.
- jsjohnst 3y agoAn important thing to also keep in mind is how much inflation changed prices over the duration. $520k in 2023 dollars is around $420k in 2017 dollars. Sure, still almost 3x more expensive, but that’s better than being 0.7x higher.
- usaar333 3y agoYour citation is for 1k A100s, not 3.5k V100s. I think it's actually ~51 days on 3.5k V100s. Just to compare the GPUs, TF32 Tensor processing went from ~125 TFlops to ~990. It then looks like they also dropped the precision to FP8, which gives you another 4x win. What's interesting is to look at how we're progressing in performance over time. In some sense, a bit slow? A V100 costs $10k at release; an H100 seems to be $40k. So we've only managed to halve the cost of a flop in 5 years. That seems.. much slower than what Moore's Law would have suggested.
- winwang 3y agoWhat about power draw? Quick google says V100 is max 300W while H100 is max 700W TDP, which makes the cost more favorable than 4x, so more like 7x less per flop. *assuming* the same electricity cost, which actually seems to have increased significantly. On a (minor) side note, it seems that $1.00 from 2018 is worth ~$1.20 in 2023. I wish more cost comparisons included inflation, because the past few years have had a lot of it.
- usaar333 3y agoGood call on power draw, but it seems dwarfed by capital cost. Assuming 3 year life (this tech gets outdated..), 300W constant consumption costs under $1k at $0.3/kwh. You need to be an order of magnitude higher in power usage before it really starts mattering. e.g. a Tesla3 consumes around 15 KW (20x the H100) while driving. That said it looks like flops/watt dropped by only 3.4x, which is also sub Moore's Law (3 years to halve power consumption)
- TOMDM 3y agoI don't think flops/$ is enough to really capture the difference here. You couldn't replicate the scale of compute this allows no matter how many V100's you had. A huge amount of cost here is embodied in networking and memory. If one were to design a chip that cared only for flop/$ without caring for all of the interconnect and memory, then the 4090 is a much fairer comparison, and even then that card isn't designed for a flop/$ optimisation.
- godelski 3y agoWith 3,584 GPUs, or 448 Nodes w/ 8GPUs/node. At 5 nodes/rack, that's 90 racks. I looked and found and old listing for Lambda's Echelon racks about $650k. So the infra cost is a little under $6 million. If you bought the $2/hr H100 instance they offer, that would cost ~$330k. Pricey, but not too bad. My bigger hope is that with cheaper compute we can see more architecture search and designs. A lot of different architectures are relatively unexplored due to computational constraints and are typically performed by smaller labs so they don't scale and it is kinda hard to compare models when we're just looking at performance benchmarks and not considering other factors. We definitely don't want big labs to railroad our research directions. Feels weird that a huge amount of NLP is based on using pretrained models and tuning them. Vision is going this way too. You basically can't get published without being SOTA so you basically have to modify an existing model or have a multi-million dollar lab and train from scratch. Really weird to expect academia to compete with big labs and really weird to not let academia take "bigger risks" and explore less popular areas. It is vital to our research path that we don't force everything onto a single track.
- m00x 3y agoH100s are now around 30k USD if you're lucky, going up to 40k on Ebay/Amazon. 3,584 GPUs at $30,000 is $104,550,000 USD only in GPUs.
- godelski 3y agoWe're talking about building a server, so usually availability is different than to a home consumer. But fwiw lambda has their Node for 330k (which gets pretty close to that price), whereas the A100's are $170k, so double my figure. Especially considering the H100 nodes are 8U instead of 4U. so you need twice as many racks.
- m00x 3y ago30k is to businesses, not consumers. Consumers don't buy H100s.
- deleted 3y ago[deleted]