4 ms·
> The cost to serve tokens is absolutely profitable today and that’s been true for at least a year. > For the data center build outs, demand for tokens is stil
by boriskourt 6mo ago
> The cost to serve tokens is absolutely profitable today and that’s been true for at least a year.
> For the data center build outs, demand for tokens is still exceeding supply.
Can you provide any numbers for this?
- bob1029 6mo agohttps://www.cerebras.ai/blog/cerebras-cs-3-vs-nvidia-dgx-b200-blackwell https://www.cerebras.ai/blog/cerebras-cs-3-vs-nvidia-dgx-b20...
- deleted 6mo ago[deleted]
- paulddraper 6mo agoAnthropic has said inference is profitable. That’s a biased source, but the math pencils. This is why switching to local open weight models saves a lot of money. (Even though it’s not apples to apples.)
- nyeah 6mo agoCan you give a few penciled numbers?
- paulddraper 6mo agoYou can rent a H100 GPU for $4/hour. [1] 300k tokens for that hour. OpenAI charges $6. Those are pessimistic assumptions. [1] https://lambda.ai/instances https://lambda.ai/instances
- drakythe 6mo ago3.99 at 8x instances, with a minimum 2 week commitment. Good luck getting 70% usage average during that time. Useful when you're running a training round and can properly gauge demand, not so great when you're offering an API.
- infecto 6mo agoIs it not a good penciled number? It helps set the directional tone that at inference cost is being covered.
- drakythe 6mo agoIt says the numbers are theoretically possible. Requiring a 66% usage to break even when 100% usage will piss off customers by invoking a queue means it’s a balancing act. “Technically correct. The best kind of correct”. So inference may technically be _capable_ of being profitable, but I have question’s about them being profitable in _practice_.
- hajile 6mo agoCan you keep that GPU 100% saturated at least 16 hours per day every day of the week? If not, you aren't breaking even.
- paulddraper 6mo agoNote this is also assuming you (1) Rent your GPUs. (2) Pay list price, no volume breaks. (3) Get only 85 tokens/sec. Realistically, frontier models would attain 200+ tokens/second amortized. Inference is extremely profitable at scale.
- aurareturn 6mo agoAssuming 80GB H100 and you inference a model that is MoE and close to the size of the 80GB VRAM, you're going to see around 10k tokens/second fully batched and saturated. An example here might be Mixtral 8x7B. You're generating about 36 million tokens/hour. Cost of Mixtral 8x7b on Open router is $0.54/M input tokens. $0.54/M output tokens. You're looking at potentially $38.88/hour return on that H100 GPU. This is probably the best case scenario. In reality, inference providers will use multiple GPUs together to run bigger, smarter models for a higher price.
- drakythe 6mo agoAnthropic also recently tweaked their usage limits to discourage use during peak hours. Why would they do that if inference was profitable?
- infecto 6mo agoDon’t confuse inference (api usage) with the consumer plan products. When people say inference is profitable they are referring to the cost to serve a token via the API. The consumer products are absolutely a question mark on profitability and as we see with most of the business and enterprise plans, going away for pure on demand use (api cost) full time.
- strangegecko 6mo agoProfitability doesn't imply infinite ability to scale. Of course they will want to prioritize their most profitable customers when they hit capacity issues.
- financltravsty 6mo agoTheir infra team is very understaffed and they are reacting to the public backlash of "no 9s?"
- paulddraper 6mo agoThose are subscription plans. They tweaked the limits/periods included in the subscription. Having higher limits for subscription plans didn't give them any more revenue.
- aurareturn 6mo agoThey do it because their demand is higher than the compute that they have available to them. Their GPUs must be melting during peak hours so they're encouraging people who move their workload to off peak hours if possible. This is the opposite of an AI bubble burst.
- ACCount37 6mo agoCheck the token prices for open weight LLMs at various independent inference providers. That gives you a very good estimate of "how much can you serve the tokens of a model of the size N for while making a profit". Now, keep in mind: Kimi K2.5 is 1T MoE. Today's frontier LLMs are in the 1T to 5T range, also MoE. Make an estimate. Compare that estimate with the actual frontier lab prices.
- lolc 6mo agoI don't think it's as easy as looking at open weight API prices. We don't know whether the operators are making a profit on all the hardware they bought. Maybe the prices we pay just cover electricity. And it's not even certain that running costs are covered by API prices: The operators may be siphoning content and subsidize from selling that. In the current volatile environment, the API prices are more of a baseline where we can assume it can't be much cheaper to operate these models.
- aurareturn 6mo agoThat doesn't make sense in this environment because everyone is compute constrained with huge backlogs they can't fulfill. If these inference providers aren't making any money, they'd simply sell their GPUs to those who are starved for compute.
- wongarsu 6mo agoI can get Kimi K2.5 inference on openrouter for about $0.5/MTok input + $2.5/MTok output, from six providers that have no moat besides efficiently selling GPU time. We can assume they are doing so at a profit (they have no incentive to do this at a loss), giving us those numbers as the cost to serve a 1T-a32b model at scale. Now we don't know the true size of any of the proprietary models, but my educated guess is that Sonnet is in about the same parameter range, just with better training and much better fine tuning and RLHF. Yet API pricing for Sonnet is $3/MTok input + $15/MTok output, exactly six times as expensive. Even Haiku is twice as expensive as Kimi K2.5. I find it difficult to believe in a world where those API prices aren't profitable. For subscription pricing it's harder to tell. We hear about those that get insane value out of their subscription, but there has to be a large mass who never reaches their limits. With company-wide rollouts there might even be a lot of subscription users who consume virtually no tokens at all.
- jerojero 6mo agoCompanies doing foundational models need to cover the cost of training which is much more expensive than training something like kimi.
- gruez 6mo ago>Companies doing foundational models need to cover the cost of training [...] But that's moving the goalposts? The original claim was on inference itself, not the whole company. > The cost to serve tokens is absolutely profitable today and that’s been true for at least a year.
- lbreakjai 6mo agoBut that's the same as thinking "This bar is selling a cocktail for $15. I could make it at home for 30 cents. They're making $14.7 dollars of profit per cocktail, the owner must be a millionaire now!" Everything is profitable if you ignore the costs.
- wongarsu 6mo agoYes. I would not consider Kimi a particularly good model relative to its size, and making a SotA model is a lot more expensive. But training costs are explicitly excluded when talking about the cost to serve tokens
- infecto 6mo agoMost/all private labs have cited inference is profitable. This was happening before the large push to scrap plans and largely charge folks the underlying api rates. Second take a look at the pricing of open models. Now certainly it’s not direct 1-1 comparison but we can use it as a baseline. Now of course folks might not be telling the truth but one of those situations where I see too many markers on the true side. For supply look at outages and growth rates at companies like openrouter. The demand is growing every week.