6 ms·
The problem space has a few aspects: 1. We're still in the "$5 airport Uber" era of LLMs. They're heavily subsidized, and everyone still complains about costs.
by FinnLobsien 4mo ago
The problem space has a few aspects:
1. We're still in the "$5 airport Uber" era of LLMs. They're heavily subsidized, and everyone still complains about costs.
2. There hasn't been a real incentive to work on cost optimization for data centers and the hardware they contain. When/if price hikes happen and send people scrambling to use other models or drastically reduce AI usage, this will suddenly need to happen.
3. We're massively overusing SOTA models. As long as you're on a subsidized subscription, you can use Claude Opus 4.8 high to write blog article meta descriptions. If you paid by token, you wouldn't do that.
4. Open models are a wildcard that could completely change the calculus.
- eru 4mo agoMostly agreed, however I'm not sure about 3: I suspect it works like gym memberships, and the companies mostly make their money from people who don't use the subscriptions all that much.
- FinnLobsien 4mo agoI think the problem is that the companies mostly don't make money, period. They may have better unit economics on underused subscriptions, but I don't see a world in which OAI/Anthropic don't heavily tighten the screws in the future. Right now it's silly to default to frontier models, but it won't bankrupt your company. I believe in the short-medium term future, we'll need to be more deliberate about model choices. In the long-term, of course, tech costs tend to plummet. Is there a future where in 15 years, my Apple Watch locally runs an Opus 4.8-class model? Maybe. And that would obviate this whole discussion.
- hparadiz 4mo ago[flagged]
- piva00 4mo agoYou can save it to your favourites, no need to comment at all if it's not going to add to the conversation.
- hparadiz 4mo agoOkay here is my adding to the conversation: The current discourse about LLMs in coding especially is based on the cheapest type of inference: text. This technology was designed for images which is a much more computationally expensive task than text. If it's already profitable to use this technology for multimedia like images and videos then using it on a text based inference for code is less then 1% as computationally expensive. Furthermore in the aggregate over time the computational expensive of text based inference precipitates negatively. In other words using it to write code will inevitably become a throwaway computational task like decompressing a jpeg. And yes decompressing jpegs would lag your 386 in the early 90s.
- piva00 4mo agoWhat? LLMs were designed for text, it's in their name "large language model". Only with specialised encoders like vision transformers they were able to process images as well but you're absolutely wrong about the original design intent. In the end you just added misinformation, just save the comment to your favourites and set a reminder to check it again in a few years like you wanted.
- hparadiz 4mo agoThe first technological breakthroughs were with face and red eye detection in 2003. Then object detection between 2008-2012. Text models didn't become useful until about 2016. Please watch the first course of Dr Fei Fei Li's lectures on the subject.
- piva00 4mo agoIf we want to keep tracing the lineage of AI we'll have to go all the way back to Markov chains from the 70s. You said LLMs were designed for images which is absolutely incorrect.
- iamacyborg 4mo agoI follow a guy called Daniel McCarthy on LinkedIn who writes a lot on CLV and that seems to be his take. Even if theoretically you get way more than you pay with subscriptions, the vast majority of people are not power users. https://danielminhmccarthy.com/ https://danielminhmccarthy.com/
- dofm 4mo agoThe vast majority of active users of ChatGPT could successfully use a model like Gemma 4 12B with agentic search if x86 hardware didn't make that so difficult. Likely even the E4B, which is really both fun and impressive. That is clearly a big component of Apple's bet, anyway.
- avadodin 4mo agoI have experimented with it and E4b is perfectly capable of being useful if you provide it with ready–to–use skills. It's still more like programming than telling a chatbot to go make you GTAVI in JavaScript and make sure the graphics are as good as the original. Maybe a safer prediction would be that most people will be fine just using hybrid agentic programs that run the models locally(probably with extra spyware). I think this is Apple's bet.
- Cthulhu_ 4mo agoI'd say that is/was their long game, but it's still very much in hype phase so there's a lot of people intensively using these models, and I don't think it's anywhere near cost efficient right now. Maybe in the long run when people get bored with it, but on the other hand people are becoming dependent on it for everyday things. We've already seen price hikes / token limits earlier this year, with suddenly some people running out of budget on the first day of the month. This will likely keep going for a while. On the other hand, costs will drop too - open models and specialized hardware, as the article notes. The long question will be whether the companies will get a return on their invested billions. I don't think they will, not with the amount of competition they're facing, and I don't think any one company or model (series) has a monopoly yet. Popularity sure, but I'm confident a competitor may appear tomorrow and people will switch.
- PunchyHamster 4mo agoTechnically yes but it's not hard to get to $20 plan caps. Till current hardware prices cool down I don't see it being easy to make money on frontier models.
- xienze 4mo ago> I suspect it works like gym memberships, and the companies mostly make their money from people who don't use the subscriptions all that much. I think it's like that, but not quite. The people who have a subscription but barely use it were probably never doing any serious work with AI in the first place. I.e., why would they get a subscription when their one or two chat questions (or, "make a picture of me as a superhero" prompts) per day can be had for free? Especially with Claude, I think people who subscribe skew very heavily towards people that can very easily make more than $20 worth of queries in a month. And then there's the not-insignificant number of people who are tokenmaxxing. It's like the gym membership model except ten percent of members are able to spend 72 hours per day at the gym while the rest spend 8 IMO.
- usef- 4mo agoBased on the people I know, they're paying because when ask they want the smartest model to be the one answering. There's still quite a difference between models.
- LUmBULtERA 4mo ago>3. We're massively overusing SOTA models. As long as you're on a subsidized subscription, you can use Claude Opus 4.8 high to write blog article meta descriptions. If you paid by token, you wouldn't do that. This idea that the subscriptions are subsidized is repeated over and over, but I've never seen any proof of this. It seems to be entirely based on the inferred API cost the subscription usage could give you, but there are a lot of assumptions needed for that to follow.
- erfgh 4mo agoThey are subsidized by the huge losses incurred by the AI companies.
- chilmers 4mo agoOnly if those losses are coming from subscriptions, instead of capex and training, which is not at all clear.
- Hamuko 4mo agoI don't understand this argument. How does it make the subscription any less subsidised if the losses are only because developing the product is just so darn expensive? Feels like arguing that it's not clear if Bugatti's losses came from selling the Veyron instead of designing and developing the Veyron.
- jstummbillig 4mo ago> We're still in the "$5 airport Uber" era of LLMs. They're heavily subsidized, and everyone still complains about costs. Inference is not exactly cheap. Based on what do you think this is "heavily subsidized" still? What would to token cost have to be, with current models, for it to not be that? What do you know that has you make such a claim?
- PunchyHamster 4mo ago> 2. There hasn't been a real incentive to work on cost optimization for data centers and the hardware they contain. When/if price hikes happen and send people scrambling to use other models or drastically reduce AI usage, this will suddenly need to happen. It's actually worse, the AI explosion just hikes hardware prices faster than capacity can catch up - and it will likely not catch up in a while because investments are both expensive, long, and might not seem all that good idea while bubble is still bubbling. The massive push frankly also made it unsustainable. If RAM didn't cost 3x and compute manufacturers would have to compete instead of selling every unit instantly at whatever price they want the frontier model tokens might've costed closer to sustainable amount
- edb_123 4mo ago> 1. We're still in the "$5 airport Uber" era of LLMs. They're heavily subsidized, and everyone still complains about costs. How does that figure look if you count in the current unprecedented LLM/AI-driven price inflation on both hardware, services and software? I don't believe we're exactly in the "$5 airport uber" era if you count that into your total.
- hirako2000 4mo agoIt's just making a parallel. We may be at the 10cent Uber. But oil and labour costs tend to go up, tokens as used today will probably cost what they cost today or less. But we won't just go to the airport, if we can go to Mars we will ask for it.
- rimliu 4mo agoIt about what you pay, not about what it costs.
- nDRDY 4mo agoTo draw a parallel - airport Ubers are still $5, but you can't buy a 2nd hand prius any more!
- edb_123 4mo agoFollowing your parallel: Except the fact that you still need a car in your life, even if you take an uber to the airport when needed :) And in this analogy you need to spend a lot more when buying a car, no matter if it's a new or 2nd hand one, following the price inflation caused by cheap Ubers. So in essence, my question is how much have those cheap Uber rides then cost you in reality, when factoring in the directly related price increases for the things you need and buy? Is it a net positive or negative at the end of the day for anyone other than the very few at the very top of the system?
- tasuki 4mo ago> Except the fact that you still need a car in your life No.
- fragmede 4mo ago> 1. We're still in the "$5 airport Uber" era of LLMs. They're heavily subsidized, and everyone still complains about costs. Do they? It's free right now at chat.com. After that it's $20/month which isn't much in the US. Three Starbucks or two meals at McDonald's will run you more than that these days.
- sfifs 4mo ago> We're still in the "$5 airport Uber" era of LLMs. They're heavily subsidized, and everyone still complains about costs. This is nonsense that AI providers want to peddle. Inference is wildly gross margin profitable - likely 90%+ gross margins. It's very easy to work out the cost structures bottoms up. All providers can drop costs to a third and still keep positive gross margins. The problems are 1. It possibly still doesn't pay out on training investment in a reasonable time frame without a massive expansion of the 90% gross margin. 2. There is no moat. As we see Mac Mini & High End GPUs stock outs and the pricing offered by DeepSeek and Qwen, the performance of Open Weight models are good enough that people can and are already shifting many inference workloads out of these 90% margin players