7 ms·
Deploying DeepSeek on 96 H100 GPUs
- 34679 1y ago"By deploying this implementation locally, it translates to a cost of $0.20/1M output tokens" Is that just the cost of electricity, or does it include the cost of the GPUs spread out over their predicted lifetime?
- dragonslayer56 1y ago” Our implementation, shown in the figure above, runs on 12 nodes in the Atlas Cloud, each equipped with 8 H100 GPUs.” Maybe the cost of renting?
- 34679 1y agoI'm confused because I wouldn't consider a cloud implementation to be local.
- deleted 1y ago[deleted]
- randomjoe2 1y agoLocal doesn't refer to "on metal" anymore to many people
- mwcz 1y ago"On metal" is muddied too. I've heard people refer to web apps running in an OCI container as being "bare metal" deployment, as opposed to AWS or whatever hosting platform. That's silly, but the idea that "local" is not the opposite of remote is even sillier.
- ffsm8 1y agoYou can run an OCI container on bare metal though. It doesn't stop being run on bare metal just because you're running in kernel namespaces, aka docker container Lots of people were advocating for running their k8s on bare metal servers to maximize the performance of their containers Now wherever that's applied to your conversation... I've no clue, too little context ( 。 ŏ ﹏ ŏ )
- okasaki 1y agoIn my opinion, if you're running k8s on bare metal, that's "k8s on bare metal" but still "<your app> on kubernetes", not "<your app> on bare metal".
- ffsm8 1y agoSorry, but then your opinion is just plain wrong Bare metal in the context of running software is a technical term with a clear meaning that hasn't become contested like "AI" or "Crypto" - and that meaning is that the software is running directly on the hardware. As k8s isn't virtualization, processes spawned by its orchestrator are still running on bare metal. It's the whole reason why containers are more efficient compared to virtual machines
- bee_rider 1y agoBare metal as in, no operating system? Does Linux really get in the way of these LLM inference engines?
- ffsm8 1y agoNo, as I said in my previous comment: bare metal as in not a virtual machine https://en.m.wikipedia.org/wiki/Bare-metal_server https://en.m.wikipedia.org/wiki/Bare-metal_server
- pessimizer 1y ago
- monsieurbanana 1y agoI missed that train
- vFunct 1y agoMy basement server really confused by all this...
- demodulation 1y agoThe one down in your Gaza tunnels?
- bee_rider 1y agoLocal doesn’t need to be “on metal,” but I’m still confused as to what they are saying. Are they running some local cloud system?
- DSingularity 1y agoI guess local for him is independent/private.
- ollybee 1y agoH100's can be $2 and hour, so $192 an hour for the full cluster. They report 22k tokens per second, so ~ 80 million an hour, thats $16 an hour at $0.2 per million. Maybe a bit more for input tokens, but it seems a long way off.
- zipy124 1y agoI think you mis-read. Thats 22k tokens per second per node, so per 8 h100's. With 12 nodes they get 264k tokens per second, or 950 million an hour. This get's you to roughly $0.2021 per million at $2 an hour.
- zipy124 1y agoThis is all costs included. Thats 22k tokens per second per node, so per 8 h100's. With 12 nodes they get 264k tokens per second, or 950 million an hour. This get's you to roughly $0.2021 per million at $2 an hour for an h100, which is what they go for on services such as runpod.io . (cheaper if not paying spot-price + volume discounts).
- adam_arthur 1y agoI'm curious as well. Depreciation and GPU failure rate over time must be considered, which I don't see mentioned in the article.
- abdellah123 1y agoWow, please edit the title to include Open-source !
- Blahah 1y agoWhy? Open source isn't in the original title
- SV_BubbleTime 1y agoAlso “open source” I feel covers for “open weights” which is not the same thing.
- adastra22 1y agoWhat does “open source” even mean when there is no source code?
- SV_BubbleTime 1y agoThere is a source, it would be the training data. There is also kind of the training code. Almost absolutely no one releases their training data.
- numpad0 1y agoThese open models are just commercial binary distributions made available at zero cost with intention to cripple opportunities for Western LLM providers to capitalize on investments. These are more like really gorgeous corporate swags than FOSS.
- badsectoracula 1y ago> intention to cripple opportunities for Western LLM providers to capitalize on investments. Western LLM providers release open weight models too (e.g. Mistral).
- echelon 1y ago
- caminanteblanco 1y agoThere was some tangentially related discussion in this post: https://news.ycombinator.com/item?id=45050415 https://news.ycombinator.com/item?id=45050415, but this cost analysis answers so many questions, and gives me a better idea of how huge the margin on inference a lot of these providers could be taking. Plus I'm sure that Google or OpenAI can get more favorable data center rates than the average Joe Scmoe. A node of 8 H100s will run you $31.40/hr on AWS, so for all 96 you're looking at $376.80/hr. With 188 million input tokens/hr and 80 million output tokens/hr, that comes out to around $2/million input tokens, and $4.70/million output tokens. This is actually a lot more than Deepseek r1's rates of $0.10-$0.60/million input and $2/million output, but I'm sure major providers are not paying AWS p5 on-demand pricing. Edit: those figures were per node, so the actual input and output prices would be divided by 12.$0.17/million input tokens, and $0.39/million output
- deleted 1y ago[deleted]
- matt-p 1y ago188M input / 80M output tokens per hour was per node I thought? Reversing out these numbers tells us that they're paying about $2/H100/Hour (or $16/hour for a 8xH100 node). Disclaimer (one of my sites) https://www.serversearcher.com/servers/gpu https://www.serversearcher.com/servers/gpu - says that a one month commit on a 8XH100 node goes for $12.91/hour. The "I'm buying the servers and putting them in COLO rate" usually works out at around $10/Hour, so there's scope here to reduce the cost by ~30% just by doing better/more committed purchasing.
- caminanteblanco 1y agoYou were definitely right, I updated the original comment. Thanks for your correction!
- caminanteblanco 1y agoOk, so the authors apparently used atlas cloud hosting, which charges $1.80 per h100/hr, which would change the overall cost to around $0.08/ million input and $0.18/million output, which seems much more in line with massive inference margins for major providers.
- s46dxc5r7tv8 1y agoSeparation of the prefill and decoding layers with sglang is quite nifty! Normally 8xH100 would barely be able to hold the 4bit quantization of the model without even considering the KV cache. One prefill node for 3 decode nodes is also fascinating, nice writeup.
- arnaudsm 1y agoInterestingly, this is 10x cheaper than the cheapest provider on OpenRouter : https://openrouter.ai/deepseek/deepseek-r1?sort=price https://openrouter.ai/deepseek/deepseek-r1?sort=price Inference is more profitable than I thought.
- brilee 1y agoFor those commenting on cost per token: This throughput assumes 100% utilizations. A bunch of things raise the cost at scale: - There are no on-demand GPUs at this scale. You have to rent them for multi-year contracts. So you have to lock in some number of GPUs for your maximum throughput (or some sufficiently high percentile), not your average throughput. Your peak throughput at west coast business hours is probably 2-3x higher than the throughput at tail hours (east coast morning, west coast evenings) - GPUs are often regionally locked due to data processing issues + latency issues. Thus, it's difficult to utilize these GPUs overnight because Asia doesn't want their data sent to the US and the US doesn't want their data sent to Asia. These two factors mean that GPU utilization comes in at 10-20%. Now, if you're a massive company that spends a lot of money on training new models, you could conceivably slot in RL inference or model training to happen in these off-peak hours, maximizing utilization. But for those companies purely specializing in inference, I would _not_ assume that these 90% margins are real. I would guess that even when it seems "10x cheaper", you're only seeing margins of 50%.
- jerrygenser 1y agoRe the overnight that's why some providers are offering there are batch tier jobs that are 50% off which return over up to 12 or 24 hours for non-interactive use cases.
- lbhdc 1y agoIf you are willing to spread your workload out over a few regions getting that many GPUs on demand can be doable. You can use something like compute classes on gcp to fallback to different machine types if you do hit stockouts. That doesn't make you impervious from stock outs, but makes it a lot more resilient. You can also use duty cycle metrics to scale down your gpu workloads to get rid of some of the slack.
- matusp 1y agoYou also need to consider that the field is moving really fast and you cannot really rely on being able to have the same margins in a year or two.
- derefr 1y ago
- ozgune 1y agoThe SGLang Team has a follow-up blog post that talks about DeepSeek inference performance on GB200 NVL72: https://lmsys.org/blog/2025-06-16-gb200-part-1/ https://lmsys.org/blog/2025-06-16-gb200-part-1/ Just in case you have $3-4M lying around somewhere for some high quality inference. :) SGLang quotes a 2.5-3.4x speedup as compared to the H100s. They also note that more optimizations are coming, but they haven't yet published a part 2 on the blog post.
- aurareturn 1y agoIsn't Blackwell optimized for FP4? This blog post runs Deepseek at fp8, which is probably the sweet spot but new models with fp4 native training and inference would be drastically faster than fp8 on blackwell.
- cootsnuck 1y agoSuper helpful to see actual examples of what it (roughly) can look like to deploy production inference workloads, and also the latest optimization efforts. I consult in this space and clients still don't fully understand how complex it can get to just "run your own LLM".
- curtisszmania 1y ago[dead]
- guerrilla 1y agoNow if only it would stop prefacing all its output with "Of course!" ;) This is why I use DS though. I think its the only ethical option due to its efficiency. I think that outweighs all other considerations at this point.
- 7thpower 1y agoThe only ethical option? Please help me understand the argument here.
- guerrilla 1y agoEverything else uses more energy for both training and inference. Reducing the energy footprint is our highest priority in this domain. It outweighs the other considerations like it being Chinese, run by a hedge fund, etc. None of that matters if we destroy our ability to live on this planet. DeepSeek is not good enough, but we need to choose it in order to encourage competition on this front specifically. It's more important that companies focus on that than spend time improving other metrics.