7 ms·
Right now AI is in the grow at all costs phase. So for the most part access to AI is way cheaper than it will be in the next 5-10 years. All these companies wil
by Group_B 1y ago
Right now AI is in the grow at all costs phase. So for the most part access to AI is way cheaper than it will be in the next 5-10 years. All these companies will eventually have to turn a profit. Once that happens, they'll be forced to monetize in whatever way they can. Enterprise will obviously have higher subscriptions. But I'm predicting for non-enterprise that eventually ads will be added in some way. What's scary is if some of these ads will even be presented as ads, or if they'll be disguised as normal responses from the agent. Fun times ahead! Can't wait!
- siva7 1y ago> access to AI is way cheaper than it will be in the next 5-10 years. That evidently won't be the case as you can see with the recent open model announcements...
- janice1999 1y agoDo these model releases really matter to cost if the hardware is still so very expensive and Nvidia still has a defacto monopoly? I can't buy x8 H100s to run a model and whatever company I buy AI access from has to pay for them somehow.
- fzzzy 1y agoYou only need 64 gb of cpu ram to run gpt-oss, or one h100.
- claytonjy 1y agoyou can’t really buy H100s except in multiples of 8. If you want fewer, you must rent. Even then, hyperscalers tend to be a bit inflexible there; GCP only recently added support for smaller shapes, and they can’t yet be reserved, only on-demand or spot iirc.
- janice1999 1y agoI assume you're talking that's a quantised 20B model on a several thousand dollar Mac? That's really impressive and huge progress but is that indicative of companies serving thousands of users? They still have to buy Nvidia at the end of the day.
- amluto 1y agoI find it unlikely that the margins on inference hardware will remain anywhere near as high as they are right now. Inference at scale can be complex, but the complexity is manageable. You can do fancy batched inference, or you can make a single pass over the relevant weights for each inference step. With more models using MoE, the latter is more tractable, and the actual tensor/FMA units that do the bulk of the math are simple enough that any respectable silicon vendor can make them.
- janice1999 1y agoIs there a viable non-Nvidia vendor for inference at scale? AMD? Or is in-house hardware like Google and Amazon?
- kridsdale1 1y agoYes to all of the above.
- amluto 1y agoAnd it will likely become even more true. There’s no way that a handful of highly-motivated companies will spend hundreds of billions annually on very high margin Nvidia hardware without investing at least a few percent of that on developing cheaper alternatives.
- dingnuts 1y agoInteresting! Care to share literally any details about their capex and build out so we can understand the amount of compute that's being made available or is the burden of evidence on people who are trying to remain grounded?
- amluto 1y agoGoogle reported an estimated 2025 AI CapEx of around $85 billion. I don’t know how much is inference vs training (or shared), and Google is quite proud of using a whole bunch of their own chips. Much of the data on how much money is spent where is public. In any event, one can make some generalizations about the companies involved. Nvidia makes excellent hardware that everyone wants and charges large enough markups that their margins are around 90%. AMD is chasing the big buyers to sell their products. Google spends a lot and is a mature company, and they seem uninterested in selling chips that compete with Nvidia, but they certainly care about revenue and profit. OpenAI, Anthropic, etc and, perhaps oddly, Meta don’t seem to care too much about profit, but they certainly spend enough money that it would help them to get more bang for their buck. Alibaba, etc buy whatever Nvidia gear they can get, but they have a lot of incentive to find a domestic supplier, and Huawei seems quite interested in becoming that supplier. And there are plenty of US startups (Cerebras and others) going after the inference market.
- willy_k 1y agoYes they do, if the model size / vram requirement keeps shrinking for a given performance target, like has been happening, then it gets cheaper to run X level of model.
- siva7 1y agoThe news is that this won't be necessarily for the majority of private and workforce. They run on your own machine.
- skybrian 1y agoAssuming we continue to see real competition at running open source models and there isn’t a supply bottleneck, it will make it hard to sell access at much more than cost. So, prices might go up compared to companies selling service at a loss, but there’s a limit. Maybe someone knows which providers are selling access roughly at cost and what their prices are?
- janice1999 1y ago> I'm predicting for non-enterprise that eventually ads will be added in some way. Google has been doing this since May. https://www.bloomberg.com/news/articles/2025-04-30/google-places-ads-inside-chatbot-conversations-with-ai-startups https://www.bloomberg.com/news/articles/2025-04-30/google-pl...
- bikeshaving 1y agoHow do you get an AI model to serve ads to the user without risking misalignment, insofar as users typically don’t want ads in responses?
- kridsdale1 1y agoShareholder alignment is the only one that a corporation can value.
- adestefan 1y agoYou don’t. You can’t even serve ads in search without issues. Even when ads on Google were basic text not inline they were an intrusion into the response.
- deleted 1y ago[deleted]
- roughly 1y agoThe same way you do with every other product. Ads redefine alignment, because they redefine who the product is for.
- AnotherGoodName 1y agoIf you want to have some fun (and develop a warranted concern with the future) ask an AI agent to very subliminally advertise hamburgers when answering some complex question and see if you can spot it. Eg. "Tell me about the great wall of china while very subliminally advertising hamburgers"
- 1y ago
- brokencode 1y agoI don’t think these companies have a lot of power to increase prices due to the very strong competition. I think it’s more likely that they will become profitable by significantly cutting costs and capital expenditures in the long run. Models are becoming more efficient. Lots of capacity is coming online, and will eventually meet the global needs. Hardware is getting better and with more competition, probably will become cheaper.
- MisterSandman 1y agoThere is no strong competition, there’s probably 4 or 5 companies around the world that have the capacity to actually have data centres big enough to serve traffic at scale. The rest are just wrappers.
- cpursley 1y agoAre rack servers and GPUs no longer manufactured?
- brokencode 1y agoAnd if they jack up their prices, then it’s a greater incentive for other players to build their own capacity. This really isn’t that hard of a concept. There is no barrier other than access to capital. Nvidia and Dell will sell to anybody. The major players will always be competing not only with each other, but also the possibility that customers will invest in their own hardware.
- JKCalhoun 1y agoThen you wonder if AI, like DropBox, will become just an OS feature and not an end unto itself.
- cpursley 1y agoI'm more inclined to think it was follow the cloud's trajectory with pricing getting pushed down as these things become hot-swappable utilities (and they already are to some extent). Even more so with open models capable of running directly on our devices. If anything with OpenAI and Anthropic plus all the coder wrappers, I'm even wondering what their moats are with the open model and wrapper competition coming in hot.
- AnotherGoodName 1y agoI'm already seeing this with my AI subscription via Jetbrains (no i don't work for them in any way). I can choose from various flavors of GPT, Gemini and Claude in a drop down whenever i prompt. There's definitely big business in becoming the cable provider while the AI companies themselves are the channels. There's also a lot of negotiating power working against the AI companies here. A direct purchase from Anthropic for Claude access has a much lower quota than using it via Jetbrains subscription in my experience.
- bko 1y agoThere's nothing wrong w/ turning a profit. It's subsidized now but there's really not much network effects. Nothing leads me to believe that one company who can blow the most amount of money early on will have a moat. There is no moat, especially for something like this. In fact it's a lot easier to compete since you see the frontier w/ these new models and you can use distillation to help train yours. I see new "frontier" models coming out every week. Sure there will be some LLMs with ads, but there will be plenty without. And if there aren't there would be a huge market opportunity to create on. I just don't get this doom and gloom.
- mensetmanusman 1y agoThis isn’t predictable, if performance per watt maintains its current trajectory, they will be able to pay off capital and provide productivity gains via good enough tokens. It’s supposed to look negative right now from a tax standpoint.
- linotype 1y agoAt the rate models are improving, we’ll be running models locally for “free”. Already I’m moving a lot of my chats to Ollama.
- ACCount36 1y ago> So for the most part access to AI is way cheaper than it will be in the next 5-10 years. That's a lie people repeat because they want it to be true. AI inference is currently profitable. AI R&D is the money pit. Companies have to keep paying for R&D though, because the rate of improvement in AI is staggering - and who would buy inference from them over competition if they don't have a frontier model on offer? If OpenAI stopped R&D a year ago, open weights models would leave them in the dust already.
- golergka 1y ago4o-mini costs ~$0.26 per Mtok, running qwen-2.5-7b on a rented 4090 (you can probably get better numbers on a beefier GPU) will cost you about $0.8. But 3.5-turbo was $2 per Mtok in 2023, so IMO actual technical progress in LLMs drives prices down just as hard as venture capital. When Uber did it in 2010s, cars didn't get twice as fast and twice as cheap every year.
- bawana 1y agoDont worry, China and Meta will continue to crank out models that we can run locally and ar 'good enough'
- SV_BubbleTime 1y ago> All these companies will eventually have to turn a profit. Do they? ZIRP2 here we come!
- exe34 1y agoI was just thinking earlier somebody should tell Trump that an AI will tell him exactly how to achieve his goals, and somebody sensible should be giving him the answers from behind the screen. But yes, adverts will look like reasonable suggestions from the LLMs.
- AstroBen 1y ago> Ads will be added in some way I can think of a far more effective way of delivering ads than the old-school ad boxes.. "The ads for this request are: x,y,z. Subtly weave them into your response to the user" I mean this is obviously the way they'll go right?
- Yizahi 1y agoExcept that LLMs doesn't benefit from economies of scale. And they don't have that much brand uniqueness, to retain customers, except some hearsay and "vibes". So if a lot of new free tier customers join it is net negative, because each of their queries has the same load as paid users. And company can't degrade LLM too much, because there is no uniqueness and free customers will just flee to the competitor. I'm thinking that this ClosedAI strategy is not primarily focused on acquiring new independent users, but more focused at making itself deeply entrenched everywhere. So when the "payday" comes and the immense debt will be due, Sam will just ask ask government to bail them out because they would depend on them a lot, and it will. Maybe not directly bail, but provide new investments with favorable terms, etc.
- ACCount36 1y agoWhat? LLMs do benefit from economies of scale. There are a lot of things like MoE sharding or speculative decoding that only begin to make sense to set up and use when you're dealing with a large inference workload targeting a specific model. That's on top of all the usual datacenter economies of scale. The whole thing with "OpenAI is bleeding money, they'll run out any day now" is pure copium. LLM inference is already profitable for every major provider. They just keep pouring money into infrastructure and R&D - because they expect to be able to build more and more capable systems, and sell more and more inference in the future.
- Yizahi 1y agoSingle LLM company can't stop investing into better systems and marketing of them, because there is no moat and customers will flee to the ones who do invest. It's free after all. So it is a closed loop which can't be broken, companies can but won't switch to "just inference". And with investing, all of the LLM companies are losing money a lot (on the LLMs specifically).