4 ms·
That's for model inference. For training OpenAI said they used around 3000 PetaFLOPS / days on the largest GPT-3 model. That translates to about 300 Nvidia A100
by exged 6y ago
That's for model inference. For training OpenAI said they used around 3000 PetaFLOPS / days on the largest GPT-3 model. That translates to about 300 Nvidia A100 GPUs if you want training to finish in a month (any slower and your researchers are not going to be able to make much progress). A system like that would cost at least $5M, probably more like $10M.
- lopmotr 6y agoIs that unit PetaFLOPS/day right? I think it should be PetaFLOPS-day which has dimensions of FLOP (total number of operations), rather than FLOP/time^2. The cost wouldn't be the cost of the hardware because it still exists afterwards. You'd have to discount it for the amount of time it was in use.
- 2sk21 6y agoSurely $10 million should be well within the spending abilities of a fair number of tech people? Many universities now have fairly large computing clusters as well.
- claudeganon 6y agoPeople give more than this to get their lackluster children into the Ivy League.