7 ms·
I've done the modeling on this a few times and I always get to a place where inference can run at 50%+ gross margins, depending mostly on GPU depreciation and h
by _sword 1y ago
I've done the modeling on this a few times and I always get to a place where inference can run at 50%+ gross margins, depending mostly on GPU depreciation and how good the host is at optimizing utilization. The challenge for the margins is whether or not you consider model training costs as part of the calculation. If model training isn't capitalized + amortized, margins are great. If they are amortized and need to be considered... yikes
- lawlessone 1y agoDoes that include legal fights and potential payouts to artists and writers whose work was used without permission? Can anyone explain why it's not allowed to compensate the creators of the data?
- ergocoder 1y agoOf course not. Those usually wouldn't be considered "margin". Another similar example is R&D and development by engineers aren't considered in margin either.
- Night_Thastus 1y agoIt's already questionable if anyone can make it profitable once you account for all the costs. Why do you think they try to squash the legal concerns so hard? If they move fast and stick their fingers in their ears, they can just steal whatever the want.
- datavirtue 1y ago[flagged]
- lawlessone 1y agowhy not?
- eikenberry 1y agoObviously because you are now allowed to download and share copyrighted works without permission or cost. At least that seems to be the precedent being set in court cases thus far.
- tick_tock_tick 1y agoBecause the law as it stands says they aren't and Congress may make no ex post facto law.
- lawlessone 1y agook, but that's just one country.
- tick_tock_tick 1y agoI mean for AI literally the only countries involved are the USA and China. I doubt you think China is going to start respecting IP rights anytime soon.
- nl 1y agoMistral is in France.
- BlindEyeHalo 1y agoWhy wouldn't you factor in training? It is not like you can train once and then have the model run for years. You need to constantly improve to keep up with the competition. The lifespan of a model is just a few months at this point.
- deleted 1y ago[deleted]
- MontyCarloHall 1y agoAs long as models continue on their current rapid improvement trajectory, retraining from scratch will be necessary to keep up with the competition. As you said, that's such a huge amount of continual CapEx that it's somewhat meaningless to consider AI companies' financial viability strictly in terms of inference costs, especially because more capable models will likely be much more expensive to train. But at some point, model improvement will saturate (perhaps it already has). At that point, model architecture could be frozen, and the only purpose of additional training would be to bake new knowledge into existing models. It's unclear if this would require retraining the model from scratch, or simply fine-tuning existing pre-trained weights on a new training corpus. If the former, AI companies are dead in the water, barring a breakthrough in dramatically reducing training costs. If the latter, assuming the cost of fine-tuning is a fraction of the cost of training from scratch, the low cost of inference does indeed make a bullish case for these companies.
- mgh95 1y ago> If the latter, assuming the cost of fine-tuning is a fraction of the cost of training from scratch, the low cost of inference does indeed make a bullish case for these companies. On the other hand, this may also turn into cost effective methods such as model distillation and spot training of large companies (similarly to Deepseek). This would erode the comparative advantage of Anthropic and OpenAI, and result in a pure value-add play for integration with data sources and features such as SSO. It isn't clear to me that a slowing of retraining will result in advantages to incumbents if model quality cannot be readily distinguished by end-users.
- trilogic 1y agoI have to disagree. The biggest cost is still energy consumption, water and maintenance. Not to mention, to keep up with the rivals in incredibly high tempo (so offering billions like Meta recently). Then the cost of hardware that is equal to Nvidia skyrocketing shares :) No one should dare to talk about profit yet. Now is time to grab the market, invest a lot and work hard, hopping for a future profit. The equation is still work on progress.
- wtallis 1y ago> The biggest cost is still energy consumption, water and maintenance. Are you saying that the operating costs for inference exceed the costs of training?
- umpalumpaaa 1y agoNo. But training an LLM is certainly very very expensive and a gamble every time you do it. I think of it a bit like a pharmaceutical company doing vaccine research…
- trilogic 1y agoThe global cost of inference in both Openai and Anthropic it exceed training cost for sure. The reason is simple: the inference cost grows with requests not with datasets. My math simplified by AI says: Suppose training GPT-like model costs = $ 10,000,000 C T =$10,000,000. Each query costs = $ 0.002 C I =$0.002. Break-even: > 10,000,000 0.002 = 5,000,000,000 inferences N> 0.002 10,000,000 =5,000,000,000inferences So after 5 billion queries, inference costs surpass the training cost. Openai claims it has 100 million users x queries = I let you judge.
- DoesntMatter22 1y agoIs that not baked into the h100 rental costs?
- tptacek 1y agoIt is.
- ProofHouse 1y agocan you share the model?
- ozgune 1y agoI agree that you could get to high margins, but I think the modeling holds only if you're an AI lab operating at scale with a setup tuned for your model(s). I think the most open study on this one is from the DeepSeek team: https://github.com/deepseek-ai/open-infra-index/blob/main/202502OpenSourceWeek/day_6_one_more_thing_deepseekV3R1_inference_system_overview.md https://github.com/deepseek-ai/open-infra-index/blob/main/20... For others, I think the picture is different. When we ran benchmarks on DeepSeek-R1 on 8x H200 SXM using vLLM, we got up to 12K total tok/s (concurrency 200, input:output ratio of 6:1). If you're spiking up 100-200K tok/s, you need a lot of GPUs for that. Then, the GPUs sit idle most of the time. I'll read the blog post in more detail, but I don't think the following assumptions hold outside of AI labs. * 100% utilization (no spikes, balanced usage between day/night or weekdays) * Input processing is free (~$0.001 per million tokens) * DeepSeek fits into H100 cards in a way that network isn't the bottleneck
- _sword 1y agoI was modeling configurations purpose-built for running specific models in specific workloads. I was trying to figure out how much of a gross margin drag some software companies could have if they hosted their own models and served them up as APIs or as integrated copilots with their other offerings
- next_xibalba 1y ago> whether or not you consider model training costs as part of the calculation Whether they flow through COGS/COR or elsewhere on the income statement, they've gotta be recognized. In which case, either you have low gross margins or low operating profit (low net income??). Right? That said, I just can't conceive of a way that training costs are not hitting gross margins. Be it IFRS/GAAP etc., training is 1) directly attributable to the production of the service sold, 2) is not SG&A, financing, or abnormal cost, and thus 3) only makes sense to match to revenue.
- lumost 1y agoI wonder how much capex risk there is in this model, depreciating the GPUs over 5 years is fine if you can guarantee utilization. Losing market share might be a death sentence for some of these firms as utilization falls.
- utyop22 1y agoWhat I hear nobody talking about is the price elasticity of demand and how this plays into the economics of the model business.
- Jaxkr 1y agoI think some of the power user demand is fairly inelastic. I’ve seen developers who are allergic to spending money happily drop $200/mo on those new Claude subscriptions.
- utyop22 1y agoYeah but if you push the price up, given that many users will cancel their subscriptions you will end up with still a tiny market segment relative to what is necessary, in revenues, to justify the valuations purported.
- lumost 1y agoIt's a tricky one, there is also a lot of push right now to use AI so developers are incentivized to drop money on subscriptions. I'd have difficulty justifying 1k/month for smaller shops - but corporations will be different. If the average engineer is just 20% more productive, then that is a 30-60k value to the company. I don't have difficulty getting to a 20% productivity gain with AI just from automating the tasks I procrastinate on or can't focus on. Likewise the ability to code a prototype overnight/over the weekend is a reasonable extension of practical working hours. The challenge I do see is that fully AI generated code bases devolve into slop pretty fast. The productivity cutoffs are much lower compared to human engineers.