11 ms·
You can't compare an API that is profitable (search) to an API that is likely a loss-leader to grab market share (hosted LLM cloud models). Sure there might no
by xxbondsxx 1y ago
You can't compare an API that is profitable (search) to an API that is likely a loss-leader to grab market share (hosted LLM cloud models).
Sure there might not be any analysis that proves that they subsidized, but you also don't have any evidence that they are profitable. All the data points we have today show that companies are spending an insane amount of capex on gaining AI dominance without the revenue to achieve profitability yet.
You're also comparing two products in very different spots in the maturity lifecycle. There's no way to justify losing money on a decade-old product that's likely declining in overall usage -- ask any MBA (as much as engineers don't like business perspectives).
(Also you can reasonably serve search queries off of CPUs with high rates of caching between queries. LLM inference essentially requires GPUs and is much harder to cache between users since any one token could make a huge difference in the output)
- xxbondsxx 1y agoFor example, Perplexity has been fudging their accounting numbers to shift COGS to R&D to make their margin appear profitable: https://thedeepdive.ca/did-perplexity-fudge-its-numbers/ https://thedeepdive.ca/did-perplexity-fudge-its-numbers/
- TZubiri 1y agoThis is addressed in the article. Giving arguments for llms being profitable as APIs.
- n4r9 1y agoOne of those arguments is: > there's not that much motive to gain API market share with unsustainably cheap prices. Any gains would be temporary, since there's no long-term lock-in, and better models are released weekly The goal may be not so much locking customers in, but outlasting other LLM providers whilst maintaining a good brand image. Once everyone starts seeing you as "the" LLM provider, costs can start going up. That's what Uber and Lyft have been trying to do (though obviously without success). Also, the prices may become more sustainable if LLM providers find ways to inject ad revenue into their products.
- unilynx 1y ago> Also, the prices may become more sustainable if LLM providers find ways to inject ad revenue into their products. I'm sure they've already found ways to do that, injecting relevant ads is just a form of RAG. But they won't risk it yet as long as they're still grabbing market share just like Google didn't run them at the start - and kept them unobtrusive until their search won.
- pr337h4m 1y agoUber and Lyft rely on network effects, which do not exist in any meaningful sense for LLM API providers.
- nitwit005 1y agoBrand is huge in every market. It's hard to get people to visit your website at all. People know about OpenAI, and look it up.
- n4r9 1y agoYeah, that's definitely a factor in the attempt to "undercut and outlast". I guess I have two defenses: firstly, network effects might not be crucial, it might be enough for there to be a small cost to changing provider; secondly, I imagine the providers are finding ways to use network effects to bolster adoption - e.g. "Find me a party date when all my friends are free, book the catering and message them with invites".
- username223 1y agoIt's addressed poorly. > First, there's not that much motive to gain API market share with unsustainably cheap prices. Any gains would be temporary, since there's no long-term lock-in, What? If someone builds something on top of your API, they're tying themselves to it, and you can slowly raise prices while keeping each increase well below the switching cost. > Second, some of those models have been released with open weights and API access is also available from third-party providers who would have no motive to subsidize inference. See above. Just like any other Cloud service, you tie clients to your API. > Third, Deepseek released actual numbers on their inference efficiency in February. Those numbers suggest that their normal R1 API pricing has about 80% margins when considering the GPU costs, though not any other serving costs. 80% margin on GPU cost? What about after paying for power, facilities, admin, support, marketing, etc.? Are GPUs really more than half the cost of this business? (EDIT: This is 80% margin on top of GPU rental, i.e. total compute cost. My bad.) Guessing about costs based on prices makes no sense at this point. OpenAI's $20/mo and $200/mo tiers have nothing to do with the cost of those services -- they're just testing price points.
- petesergeant 1y ago> What? If someone builds something on top of your API, they're tying themselves to it, and you can slowly raise prices while keeping each increase well below the switching cost. Have you used any of these APIs? There's very little lock-in for inference. This isn't like setting up all your automation on S3, if you use the right library it's changing a config file.
- jsnell 1y ago> What? If someone builds something on top of your API, they're tying themselves to it, and you can slowly raise prices while keeping each increase well below the switching cost. That's not really how the LLM API market works. The interfaces themselves are pretty trivial and have no real lock-in value, and there's plenty of adapters around anyway. (Often first-party, e.g. both Anthropic and Google provide OpenAI-compatible APIs). There might initially have been theories that you could not easily move to a different model, creating lock-in, but in practice LLMs are so flexible and forgiving about the inputs that a different model can be just dropped in an work without any model-specific changes. > 80% margin on GPU cost? What about after paying for power, facilities The market price of renting that compute on the market. That's fully loaded, so would include a) pro-rated recouping the capital cost of the GPUs, b) the power, cooling, datacenter buildings, etc, c) the hosting provider's margin. > admin, support, marketing, etc.? Are GPUs really more than half the cost of this business? Pretty likely! In OpenAI's leaked 2024 financial plan the compute costs were like 75% of their projected costs.
- pama 1y agoPlease read the DeepSeek analysis of their API service (linked in this article): they have 500% profit margin and they are cheaper than any of the US companies serving the same model. It is conceivable that the API service of OpenAI or Anthropic have much higher profit margins yet. (GPUs are generally much more cost effective and energy efficient than CPU if the solution maps to both architectures. Anthropic certainly caches the KV-cache of their 24k token system prompt.)
- iamnotagenius 1y agoWith all due respect to Deepseek, I would take their numbers with grain of salt, as they might as well be politically motivated.
- jarym 1y agoAny more politically motivated than a model from anywhere else?
- WithinReason 1y agois that better or worse than commercially motivated?
- leeoniya 1y agocommercial motivatation needs to show eventual profit to be sustainable, while political does not. though at the outset (pre-profit / private) it's hard to say there's much difference.
- bee_rider 1y ago> though at the outset (pre-profit / private) it's hard to say there's much difference. I think this is the tough part, we’re at the outset still. Also, a political investment could could be sustainable, in the sense that China might decide they are fine running Deepseek at a loss indefinitely, if that’s what’s going on (hypothetically. Actually I have never seen any evidence to suggest Deepseek is subsidized, although I haven’t gone looking).
- JimDabell 1y ago> you also don't have any evidence that they are profitable. Sure we do. Go to AWS or any other hosting provider and pay them for inference. You think AWS are going to subsidise your usage of somebody else’s models indefinitely? > All the data points we have today show that companies are spending an insane amount of capex on gaining AI dominance without the revenue to achieve profitability yet. Yes, capex not opex. The cost of running inference is opex.
- rco8786 1y agoAWS isn’t doing the training on those models.
- JimDabell 1y agoOpenAI spends less on training than inference, so the worst case scenario is less than double the cost after factoring in training. Inference is still cheap.
- rco8786 1y agoInference is cheap. Training is cheaper. Then where's all the money going? OpenAI is reporting heavy losses, but you're saying the unit economics of inference are all good. What are they spending money on?
- mediaman 1y agoSalary, mostly. It's useful to separate out the GPU cost of training from the salary cost of the people who design the training systems. They are expensive. That does not mean, however, that inference is unprofitable. The unit economics of inference can be profitable even while the personnel costs of training next-generation models are extraordinary.
- jsnell 1y agoTheir spending is not a problem. It's quite low for a top-tier hard tech company that's also running a consumer service with 500M active users. They are making a loss because 95% of their users are on free accounts, and for now they're choosing not to monetize those users in any way (e.g. ads).
- Workaccount2 1y agoJust wait till there are ads for free users, which is going to happen. Depending on how insidious these ads are, they could be extremely profitable too, like recommending products and services directly in context.
- Sevii 1y agoThey could dynamically update the system prompt with ad content on a per request basis. Lots of options.
- handfuloflight 1y agoWhy do you equate contextual with insidious?
- Workaccount2 1y agoBecause then the AI isn't working for you anymore, it's working for the advertisers. Which isn't necessarily bad, but we can be pretty confident that the AI will not be upfront about this, and instead try to act like it's working for you.
- handfuloflight 1y agoIf the advertising is contextually relevant, how is it working against you?
- pigeons 1y agoJust being contextually relevant doesn't mean its in your interests as opposed to in the interests of the advertised or that the levers are transparent.
- handfuloflight 1y agoAre you assuming all commercial relationships are adversarial? Why can't advertisers and those advertised to have aligned interests? What transparent levers do non-advertised results have, how do you know search rankings don't have hidden commercial incentives? Why trust undisclosed bias over disclosed relationships? Isn't transparency about incentives better than pretending they don't exist?
- lumost 1y agoWe don’t know what the marginal cost of inference is yet however. So far, users are demonstrating that they are willing to pay more for LLMs than traditional web experiences. At the same time, cards have gotten >8x more efficient over the last 3 years, inference engines >10x more efficient and the raw models are at least treading water if not becoming more efficient. It’s likely that we’ll lose another 10-100x off the cost of inference in the next 2 years.
- Palmik 1y ago> API that is likely a loss-leader to grab market share (hosted LLM cloud models). I don't think so, not anymore. If you look at API providers that host open-source models, you will see that they have very healthy margin between their API cost and inference hardware cost (this is, of course, not the only cost) [1]. And that does not take into account any proprietary inference optimizations they have. As for closed-model API providers like OpenAI and Anthropic, you can make an educated guess based on the not-so-secret information about their model sizes. As far as I know, Anthropic has extremely good margins between API cost and inference hardware cost. [1]: This is something you can verify yourself if you know what it costs to run those models in production at scale, hardware wise. Even assuming use of off-the-shelf software, they are doing well.
- noodletheworld 1y agoI don’t completely disagree, but “assertion one” [1] [1] ~ you can obviously verify this yourself by doing it yourself and seeing how expensive it is. …is an enormously weak argument. You suppose. You guess. We guess. Let’s be honest, you can just stop at: > I don’t think so. Fair. I don’t either; but that’s about all we can really get at the moment afaik.
- vslira 1y agohe's not wrong, if you can run a open weights model in any cloud, you can very straightforwardly estimate the cost of running the model. considering that these providers either use long-term contracts or maybe even buy their own hardware, this theoretical cloud deployment is itself an overestimate of the costs
- noodletheworld 1y ago…and its perfectly legit to run that, write the numbers down and link to it. But: A) it makes absolutely no difference to the fact you have no idea what the big LLM providers are actually doing. B) Just asserting some random thing and saying “anyone competent can verify this themselves” is a weak argument. Youre saying youve done the research, but failing to provide any evidence you actual have If youve crunched the numbers then man up and post them. If not, then stop at “I think…” “This is based on my experience running production workloads…” is a nice way of saying “I dont have any data to backup what Im saying”. If you did, you could just link to it. …by not posting data you make your argument non-falisifyable. It is just an oppinion.
- ozim 1y agoI think you can make an educated guess if you check local model performance, prices of energy and hardware and price of the subscriptions. Best part is you can make perplexity research task out of it
- jstummbillig 1y agoThere is also a lot of different models at a lot of different price points (and LLMs are fairly hard to compare to begin with). In this theory of a likely loss-leader, must we assume that all of them, from all companies, are priced below cost...? If so, that seems like a fairly wild claim. What's Step 2 for all of these companies to get ahead of this, given how model development currently works? I think the far more reasonable assumption is: It's profitable enough to not get super nervous about the existence of your company. You have to build very costly models and build insanely costly infrastructure. Running all of that at a loss without an obvious next step, because ALL of them are pricing to not even make money at inference, seems to require a lot of weird ideas about how companies are run.
- otterley 1y agoWe’ve seen this pattern before. This happened in the 1990s during the original dot-com boom. Investors gamble, everything is subsidized, most companies fail, and the ones left standing then raise prices.
- dietr1ch 1y agoI don't think it's that wild. Hardware will improve together with performance, but once the market stops expanding and behaviour gets stagnant the market shares will solidify, so you better aim to have a large portion to make the scale together with the improvements help reach profitability.
- ddp26 1y agoI analyzed OpenAI API profitability in summer 2024 and found inference for gpt-4 class models likely pretty profitable, ~50% gross margins (ignoring capex for training models): https://futuresearch.ai/openai-api-profit https://futuresearch.ai/openai-api-profit
- otterley 1y agoThat’s a little like saying you can compute the profitability of the energy market by looking only at the margins of gas stations. You can’t exclude all the outlays on actually acquiring the product to sell.
- lazide 1y agoSure - but is there any doubt in that example that gas stations are making a profit? And unlike gasoline, once models are trained there is no significant ongoing production cost.
- otterley 1y agoModels aren't static. In order for them to remain relevant, they have to be constantly retrained with new data. Plus there's a model arms race going on and which will probably continue for the foreseeable future.
- lazide 1y agoFair point - though various distilling and retraining tricks do reduce the cost quite a bit. It’s not like everyone is doing all the work they had to do from scratch, every time.
- int_19h 1y agoThe problem with this theory in general is that, given the sheer number of cloud inference providers (most of which are hosting third party models), it would be exceedingly strange if not only all of them are engaging in this same tactic, but apparently all of them have the same financial capacity to do so.
- raincole 1y ago> an API that is likely a loss-leader to grab market share (hosted LLM cloud models) Everyone just repeats this but I never buy it. There is literally a service that allows you to switch models and service providers seamlessly (openrouter). There is just no lock-in. It doesn't make any financial sense to "grab market share". If you sell something with UI, like ChatGPT (the web interface) or Cursor, sure. But selling API at a loss is peak stupidity and even VCs can see that.
- julianeon 1y agoThis is also the argument of the guy in the article, fyi (it's not a loss leader, no reason for it to be).
- mupuff1234 1y agoExcept they most likely do have a plan to make it harder to switch.
- raincole 1y agoYeah, sure, please elaborate on how providers such as Fireworks, DeepInfra, Chutes are going to "make it harder to switch."
- mupuff1234 1y agoI'm talking about openAI, anthropic, Google, etc. They'll offer consumer and enterprise integrations that will only work with their models.
- hedayet 1y agoyes. And they will try both carrots and sticks. The carrots are already visible - think abstractions like "projects" in ChatGPT.
- DarmokJalad1701 1y agoWho is "they"? It makes no sense for Openrouter to allow providers that do not conform to the API. They profit from the commission from the fees and not providing inference.
- cush 1y ago> You can't compare an API that is profitable (search) to an API that is likely a loss-leader to grab market share (hosted LLM cloud models). Regardless of maturity lifecycle, by definition loss-leaders are cheap. If I go to the grocery store and milk is $1, I don't think I'm being swindled. I know it's a loss-leader and I buy it because it's cheap. We are currently in the early-Netflix-massive-library-for-five-dollars-a-month era of LLMs and I'm here for it. Take all you can grab right now because prices will 100x over the next two years.
- jackdeansmith 1y agoWant to bet? I'll give you 5:1 odds that tokens from a model with some specific benchmark performance (we can sort out the specific benchmarks or basket of benchmarks if you want to bet) will be cheaper two years from now.
- cush 1y agoSure. To clarify, I'm not asserting that in two years today's 4o will be 100x more expensive, but the sum of many core offerings from various companies will be. It won't be unheard of for people to spend $2k-$10k/yr between many AI services
- chipsrafferty 1y agoIt's not unheard of for people to spend $13,000 on a pure silver frying pan. Is it common? No.
- cush 1y agoLike say 10% of people or more
- jackdeansmith 1y agoI misunderstood your comment then, my assertion is that models which have the same capabilities as current models will be cheaper in the future. I have no doubts that models with more capabilities will be more expensive.
- luqtas 1y ago+ not considering the amount of copyright violations on the training weights, if it was easy and cheap for the masses to use the judiciary system, maybe this technology would be way behind of what it's "capable"
- jongjong 1y agoYep spot on. Price does not equate cost. Especially in our current economy where profit has been artificially made a non-factor. To know the cost, you'd have to look at hardware resource usage per query. Given that recent models have over a trillion parameters, you need a huge amount of memory and CPU to process a query to get the electrons to traverse all these thousands of billions of ANN nodes and/or weights. Ultimately, it may turn out that dumber models may be more economically efficient than smarter models once you ignore the investment subsidy factor. Maybe, given the current state of AI, the economically efficient situation is to have lots of dumb LLMs to solve small, well-defined problems and leave the really difficult problems to humans. Current approach, looking at pricing is assuming another AI breakthrough is just around the corner.