4 ms·
Full title & subtitle: DeepSeek might not be as disruptive as claimed, firm reportedly has 50,000 Nvidia GPUs and spent $1.6 billion on buildouts The fabled $
by sxp 2y ago
Full title & subtitle:
DeepSeek might not be as disruptive as claimed, firm reportedly has 50,000 Nvidia GPUs and spent $1.6 billion on buildouts
The fabled $6 million was just a portion of the total training cost.
- coldtea 2y agoThey also run it as a service though? How much of that hardware costs was training vs infrastructure for the public service? Besides, it's FOSS, isn't? Meaning anybody can see how much it takes to run and how much it takes to train? >A recent claim that DeepSeek trained its latest model for just $6 million has fueled much of the hype. However, this figure refers only to a portion of the total training cost— specifically, the GPU time required for pre-training. It does not account for research, model refinement, data processing, or overall infrastructure expenses. So the $6 million is correct - and the rest is irrelevant additional costs. Yes, they'd pay for research. They'll pay for data processing. They'll have overal infrastructure expenses. Duh!
- stonogo 2y agoIt is not FOSS. The LLM industry has repurposed "open source" to mean "you can run the model yourself." They've released the model, but it does not meet the 'four freedoms' standard: https://github.com/deepseek-ai/DeepSeek-V3/blob/main/LICENSE-MODEL https://github.com/deepseek-ai/DeepSeek-V3/blob/main/LICENSE...
- coldtea 2y agoThe code is MIT however.
- throwaway314155 2y agoIn fairness, you can still reverse engineer the training procedure yourself from their paper and get a close approximation of training cost using some open/synthetic datasets. You might think it's incredibly complicated to do this without source code, but the pre-training portion of the training is something you can grab from other projects or re-derive yourself pretty easily (and will account for the bulk of training time/cost). Not defending it, but it's a much better situation than some of the non-commercial licensing from e.g. Meta, Stability.
- ericye16 2y agoWithout the actual training corpus, you can't know how much it takes to train. They could have trained on twice as many tokens for example (not saying they did!)
- coldtea 2y agoYou could however train it yourself and see what it takes to get the same cognitive performance.
- drakenot 2y agoYes. And the original announcement and paper called out that this ~6 million was just for training the R1 model. It specifically mentioned that this didn't include any other costs, including research, previous models, etc. Just raw training costs for the latest model. There is Hugging Face's "Open-R1" project which is attempting to replicate the DeepSeek-R1 model given the innovations outlined in the paper(s). They need to replicate the training code, the datasets, etc. This part of the system wasn't published. And you are right that with the open weights of the R1 model, it should be pretty easy to infer how much it costs to _run_ as well.