8 ms·
AI's economics don't make sense
Related: AI's biggest critic has lost the plot - https://news.ycombinator.com/item?id=47934353 https://news.ycombinator.com/item?id=47934353
- wonderwhyer 5mo agoYeah. And weird pricing seems like it's winding down. It's interesting to compare it to electricity. Basically Anthropic was selling a flat fee electricity subscription, and when someone started connecting expensive washing machines (OpenClaw) to their subscriptions, instead of changing the pricing model, they banned washing machines... I wonder if we will get to "electricity" style pricing for AI. What makes electricity predictable is relatively constant average usage over time + price is manageable. I'm just not buying electrical house heating and manage my electricity spending within some bounds. With AI the problem is that we are only now getting to useful AI, and for now it's still too expensive to be useful, so they subsidize until they can stabilize at "cheap enough and smart enough" level. But it feels like that's still 2 years away while they are stopping to subsidize now. Will be interesting.
- linkregister 5mo agoOpenClaw was never banned from the Claude API, only flat-fee plans.
- gruez 5mo ago>Basically Anthropic was selling a flat fee electricity subscription No? It was flat, but with ambiguously stated limits (eg. 5x, 10x 20x). They were discriminating on how the "electricity" was used, but that's not that much different than how power companies have different rates for residential users vs industrial users.
- ethin 5mo agoEven now they are insanely ambiguous with respect to their usage limits. They don't from what I know openly disclose them anywhere, so them saying "5x increase" is utterly meaningless, alongside "20x" or "10x" or whatnot, because we don't know what "x" is.
- swader999 5mo agoThe Uber subscription analogy works well too.
- jcgrillo 5mo agoThe finding out phase has begun.
- asah 5mo agomeh - by this logic, every new tech and startup ever is a "scam" The truth is that the AI companies are gambling that inference cost will continue following a hyper version of Moore's Law, e.g. Google TurboQuant. The countervailing thesis is that frontier models are consuming more and more compute. The deepest truth: you often don't need a frontier model to get commercially acceptable results from AI. Thus, bring on the true pricing! and I'll just switch models to something financially sustainable.
- swader999 5mo agoWe work comes to mind. The math is fairly easy if we know what a company like OpenAI's datacenter commitments are, what their sub and token revenue is right now and what their operation costs are. This is very basic and if you had that info you would know exactly if we are in bubble or not. Waiting for the S-1's...
- wood_spirit 5mo agoThe general problem the average user has with a metered instead of provisioned billing model for computer services is you can’t easily control for cost overruns. From the old days customers getting stung for hosting costs when slashdotted or DOSed, to last decades microservice shock horror of the CI retry loop that burns money overnight to today’s AI that you basically have no idea how efficient the AI will be while it ponders your question, you are just setting yourself up for disappointment and cost overruns and a feeling that you’re not getting the value for money you got last week etc.
- gruez 5mo ago>The general problem the average user has with a metered instead of provisioned billing model for computer services is you can’t easily control for cost overruns. Is this an actual issue aside from people letting their autonomous agents run overnight?
- wood_spirit 5mo agoI can speak of myself. Sometimes my session starts out well and I get the AI to cruise to 80%. But then gains after that seem impossible and what was built steadily unravels and then I get the compacting conversation message and realise that I’ve just spent a lot of money on nothing.
- deleted 5mo ago[deleted]
- lbrito 5mo ago>At some point, the incredible, toxic burn-rate of generative AI is going to catch up with them, which in turn will lead to price increases, or companies releasing new products and features with wildly onerous rates (..) that will make even stalwart enterprise customers with budget to burn unable to justify the expense. I pray this happens soon, but I feel I've been hearing some version of it for a while.
- ambicapter 5mo agoBig ships take a while to turn.
- ToucanLoucan 5mo agoThe only reason it hasn't is the sheer amount of credit being thrown at this tech. Both that and the valuations of the firms in question is stratospherically over-hyped and over-valued. This tech has uses. It has quite a lot of them in fact. However there is no usage of ChatGPT or Claude that makes OpenAI or Anthropic worth anything fucking close to what they're valued at right now, and both firms are scrambling to figure out how to get down from the top of the AI house of cards without detonating in the process. Meanwhile DeepSeek is coming out with more capable models that run on far less onerous hardware and with far less compute requirements that does basically exactly what the vast majority of users actually want it to do. This is going to be a financial bloodbath. Not for anyone actually responsible for it, of course, they'll be fine. It'll be everyone else getting soaked which is the only reason I give two shits.
- joshjob42 5mo agoThere's a few major problems with the article. The most obvious is that frontier labs are not charging remotely close to the cost of tokens; afaik most estimate north of 80% profit margins. As a reference, providers are profitably providing Kimi K2.6 for $4/1Mtok out. Is that as good as Opus? No, but it's probably at least Sonnet level, so that's ~4x cheaper than Sonnet while still being profitable to serve on the margin. So you aren't plausibly getting into actual subsidization territory until you're over 5:1 sub to nameplate token costs. How many tokens can you realistically burn through in one chat session? Opus and many other frontier models do maybe 60tok/s, less 250k/hr out. In you can use more, but in most cases cache is 5-10:1 cheaper than new input. Say you average 500ktok in, 90% cache, per request. That amounts to 100-150ktok in new input-equivalent costs, which in most cases is ~20-30ktok in output-equivalent costs. Do a request every minute, that's a total of about 1.5-2Mtok/hr. At API prices that's $50/hr for Opus, but really it probably only costs Anthropic $10/hr to serve that. That said, even if a developer is burning $50/hr, many, many employees at large companies cost more than $100k/yr to employ all costs considered, so making them say 20-30% more productive can easily make that worth it for most. If the labs shave their margins ultimately to more like 20-30%, you'd have ~$15/hr in costs to use the services, and nearly every white collar job is way over 30k/yr to employ. If your salary is 80k, you probably cost the company 200k all in, so making you 15% more productive offsets the $15/hr cost. So first party providers are not in a horrifying position or anything from a subsidization standpoint. The people in bad shape are Cursor and Perplexity, who don't have frontier models and are dependent on the open source community, which is typicly 6-12 months behind the frontier. They have to pay full freight API costs at 80% margin for the big boys to serve their harnesses, which is indeed untenable, and they'll have to either force users to use open source models and/or in house models they can serve at-cost or they will have to charge vastly more. Gemini, Claude, and ChatGPT first-party services like Antigravity, Codex, and Claude Code are not in serious trouble though.
- ToucanLoucan 5mo ago> That said, even if a developer is burning $50/hr, many, many employees at large companies cost more than $100k/yr to employ all costs considered, so making them say 20-30% more productive can easily make that worth it for most. If the labs shave their margins ultimately to more like 20-30%, you'd have ~$15/hr in costs to use the services, and nearly every white collar job is way over 30k/yr to employ. If your salary is 80k, you probably cost the company 200k all in, so making you 15% more productive offsets the $15/hr cost. Nobody including the connected article is making the argument that this cannot be profitable ever. People are saying "there is no way this admittedly quite interesting tool is going to be able to make back all of this money" and I think they are completely right to say that. You can absolutely make money with this stuff, just not at this scale. The buildout for this shit has been certifiably crazy and a number of the involved firms are overleveraged for tens and even hundreds of billions of dollars. How in the sweet fuck are you paying that off, plus giving investors dividends, selling this at $15/hour/user??? That math does not math. A quick google says there are between 1.5 and 4.4 million developers in the US alone, let's say it's 5 million, to be generous, and each of them is subbed to this for 8 hours per day, continuously. That's 600 million per year in revenue. If you took ALL that revenue, and put it towards paying down this debt, not leaving any for employee salaries, upkeep, ongoing development, it would take DECADES to pay down what OpenAI already owes. And yes I'm sticking directly to code, because that's the only thing I've seen it be really good at. Are we really proposing that every knowledge worker on earth and every manager of such workers is going to have an autonomous agent running all the time!? To do what, make sure they don't have to read or write email? Which even just that example is bringing in a fucking mess of legal, compliance, and security violations because LLMs are not intelligent and are not capable of being properly secured. Like I'm sorry, I cannot take this industry seriously when even the most basic back-of-napkin math is saying, nay, screaming from the rooftops that they are FUCKED.
- cheeseblubber 5mo agoIt make sense if you account for cost of intelligence getting cheaper every year. Most of the models per unit of intelligence is getting far cheaper. We get better hardware, architecture, training techniques, inference optimizations and caching. All those improvements add up. In in early 2022 you were getting 10x cheaper annually now is closer to 2x - 5x cheaper annually. The cost is still dropping where as Uber can only get the cost down by so much.
- mkesper 5mo agoBetter hardware would have to be bought with additional money. And no one can forecast reliably how much optimization is left in the game.
- cheeseblubber 5mo agoMy problem with the article is that they don't even mention this fact. The metaphors with Uber often is brought up but it breaks down at cost optimization. It also wouldn't be fair to say we are at the peak efficiency of LLMs and that there wouldn't be any improvements left.
- iooi 5mo agoThe entire basis of this article is that generating tokens is a variable cost and that that cost will not decrease over time. > On an economic basis, a monthly subscription only makes sense with relatively static costs. Running a data center is a fixed expense. Whether or not people use that data center to it's capacity doesn't change how much the operator pays (electricity use factors into this, since a GPU running at 100% will use more watts than an idle one, but it doesn't move the needle much on other fixed and variable costs of a data center). > They also assumed, I imagine, that the cost of tokens would come down over time, versus what actually happened — while prices for some models might have come down, newer “reasoning” models burn way more tokens, which means the cost of inference has, somehow, gotten higher over time. This is backwards. When the cost of something goes down, people use it more. This is basic supply and demand. Inference has gotten cheaper already, and will continue to do so. Companies subsidizing costs for growth happens all the time. Yes, switching to usage-based pricing instead of subscriptions sucks for customers, but enterprises will continue to pay.
- xnx 5mo ago> it doesn't move the needle much on other fixed and variable costs of a data center I wonder what the rough costs of a data center look like over the lifetime of one GPU generation? 10% building 60% GPU 30% power I haven't gone looking for that information, but I haven't run across it either.
- Marciplan 5mo agoI am a paying subscriber to Ed Zitron and I enjoy his writing a lot. He should at some point admit that not everything is bullshit and there is definitely a business model to it. It is fun to read, though
- xnx 5mo agoIt's good to have contrarian viewpoints, but Ed Zitron is so blinded by his AI hate that his articles should be treated not just with skepticism, but heavy suspicion.
- mediaman 5mo agoHe has a fun writing style but has so many willful errors, and is so committed to one point of view regardless of the facts, that his writing seems kind of worthless. I soured on him when he could not calculate cumulative revenue on an exponential curve, ignored everyone who showed him how to calculate it, and then kept writing that Anthropic’s revenue numbers are fake based on his inability to do math. It’s too bad because any heavily hyped industry needs good critics (think Ida Tarbell to Rockefeller) but they should be honest critics, and he’s not, which really undermines not only his but others’ criticism of the industry.
- putzdown 5mo agoThe moves from “the subscription model for AI isn’t working given these parameters” to “a subscription model for AI can never work” to “the model was deliberately deceptive” to “it’s a fucking ripoff” is not logical. AI companies are feeling the need to get hold of spiraling costs by increasing prices and limitations. Inference hasn’t gotten cheap enough fast enough, and for some reason they feel they can’t wait longer. That doesn’t mean a subscription service can’t work: only that it will be expensive, maybe vastly so, and will need tiers based on usage with some fluidity for users to move between tiers in a given month. The model is something like HP’s “instant ink” service. Sure, there’s a question whether the moves companies are making now are worth the cost in the eyes of customers. But that’s a question of economics and timing, not a fundamental blow to monthly subscriptions as a model. The article doesn’t deal with these considerations fairly. It’s too much in the direction of a rant, with conspiracy theories thrown in.
- christkv 5mo agoI'm just flabbergasted at the massive inefficient usage of tokens. What are people doing to spend 500 usd/day in tokens. I just don't understand what you could possibly be doing that would be not complete spagetti at the end if you run something in an autoloop.
- xnx 5mo ago> What are people doing to spend 500 usd/day in tokens 1) They're lying 2) Status signalling
- christkv 5mo agoThere is status in showing your inefficiency ?
- mrguyorama 5mo agoThat's almost all status signalling ever is.
- xnx 5mo agoA $500 Gucci belt doesn't hold up your pants any better.
- doctoboggan 5mo agoUsing Claude code with Opus 4.7 and xhigh effort for a few hours will definitely cost hundreds of usd. I am not sure if you would call claude code "an auto loop", but you don't need to be running something crazy like gas town to spend a lot of tokens with Claude.
- intended 5mo agoIt looks like a “People respond to incentives (prices)” situation. If something is cheaper than alternatives, spending patterns change. People subsidize corn or power and so consumers alter behavior to take advantage of those prices.
- georgeburdell 5mo ago
- milesvp 5mo agoReading this piece, I'm reminded of a podcast I heard some years ago where they were interviewing an early google marketing employee who was talking about the economics of google search. They said they'd done some surveys and concluded that they determined that the average user would get something like $20/year of value, and so that was the most they could realistically charge for search. Meanwhile, they could make something like $500/user in Q4 alone for advertising. So, of course, advertising. I just don't think that LLM business models can survive the allure of advertising dollars, any more than Search could, or TV, or Radio, or Movies. Ignoring the talk of copilot putting ads into pull requests, there is just no way that publicly hosted LLMs will not end up inserting ads into the output. This looks like what I remember. https://freakonomics.com/podcast/is-google-getting-worse/ https://freakonomics.com/podcast/is-google-getting-worse/
- swader999 5mo agoThe output won't be read by humans (and increasingly this is the case in my own use) so I don't see how that works. If the output itself will be directed by the highest bidder, that doesn't work. Or if the output influences the agent's direction, that doesn't work either.
- meheleventyone 5mo agoThey could make it work like rewarded video ads in mobile games. Block progress until you watch the ad. Then as dutiful engineers people can consume ads to support the business and avoid being laid off. More seriously for software engineering it’ll just cost a lot.
- gizajob 5mo agoStallman is going to be overjoyed when all the class and variable names in open source repositories have been reformatted to say EnjoyCocaCola and year_of_the_trucks_medicated_pad etc
- leecommamichael 5mo agoPlease don't give them ideas. :(
- Ritewut 5mo agoIt makes sense when you realize the goal is not the consumer but large gov and enterprise contracts.
- throwawayajner 5mo agoZitron misunderstands the economics of models. Inference costs have dropped 99% in less than 2 years. Models are being commoditized faster than any technology in history. A $20 subscription 2 years ago is not providing the same level of intelligence you're getting today. Every major lab knows open source models are 6 months behind (See Google's "We have no moat") and none of them plan to make money on inference. Companies are subsidizing users to create moats that persist when models are essentially free for most everyday use.
- pmdr 5mo ago> A $20 subscription 2 years ago is not providing the same level of intelligence you're getting today. That subscription was then and is now likely still subsidized.
- davikr 5mo agoFor all we know, there could be 10 people paying for a ChatGPT subscription and not using it enough to subsidize 1 power user _and_ still have money left for profit.
- pmdr 5mo agoOh they'd be sure to let us know if that were the case.
- warkdarrior 5mo agoWhy would the AI companies advertise that most of their users do not use their subscription in full??
- deleted 5mo ago[deleted]
- mNovak 5mo agoDo we know the breakdown of revenue from API vs subscriptions for OAI/Anthropic? That seems very relevant, since this entire article seems to be on the premise that users are only willing to pay for a subsidized subscription and would never pay the 'true' token cost. The internet seems to be saying that 70%+ of Anthropic revenue is per-token metered API, which would largely invalidate the article, but I can't find a solid source.
- swader999 5mo agoI don't think these companies will give this information up until their hand is forced with an S-1 when they want to IPO. So stay tuned...
- feverzsj 5mo agoIt makes perfect sense, if you treat it as a Ponzi scheme. [0]: https://www.wheresyoured.at/why-are-we-still-doing-this/ https://www.wheresyoured.at/why-are-we-still-doing-this/
- ameliaquining 5mo agoAs it happens, published just this morning is an article from Kelsey Piper that explains in some detail what's wrong with Zitron's takes: https://www.theargumentmag.com/p/ais-biggest-critic-has-lost-the-plot https://www.theargumentmag.com/p/ais-biggest-critic-has-lost...
- Darwins_Toffees 5mo ago- Reproduce academic papers - Put coding projects online for me so I can share them with friends - Determine which books in a set are missing from the school library and find where they’re cheapest online - Figure out which soccer club the team I see practicing at the local rec center belongs to and how to register my son - Design a bunch of robot-themed handwriting activities for a kindergartner who needs to practice making his uppercase and lowercase letters distinct I'm sorry but telling me that this is what AI can do is a sad state of affairs. Like this is google level stuff.
- 1attice 5mo agoI read that and I found it unconvincing. KP is correct that EZ is, by now, emotionally and perhaps ideologically fixated on AI's approaching reckoning, but that's KP psychologizing about Ed's inner states, which is neither fruitful nor relevant to consider when confronting a reasoned argument (or, in Ed's case, several.) EZ might have incautiously and incorrectly called the peak several times, but his newsletter is nearly always stacked with citations and insights that, at least to my cursory but frequent inspection, pan out. His argument(s) have evolved over time, but what of it? That just shows he's not the dogmatist the author wants him to be. Discourse evolves, get over it. 2026 Zitron has a good sense of the scale at which AI is requiring enormous financial complexity and volume to realize, and his basic point is that it isn't sustainable in the medium term. He is self-evidently correct.
- matchagaucho 5mo agoSame debate as the dot-com era. Customer: “I don’t want to pay more than $100/mo for my website” Developer: “What are your goals?” Customer: “1M daily visits, 1,000 monthly signups.” And we've spent the past 25 years offering serverless compute, auto-scaling, pay-as-you-go for AWS and Internet infrastructure. And the economics are still a hard sell.
- threepts 5mo agoI thought this burning of cash was all an excuse for the exponential growth we saw in the last 6 years. They went from GPT 2 a text only, goldfish-esque memory at a 8th grade reading level to what we have today, GPT 5, multimodality + a token window encompassing a enclyopedia and a Doctorate/Masters level of mastery in major subjects. The economics are probably betting on this exponential growth to continue, which if it fails, the cash would burn.
- Glyptodon 5mo agoI think there's another route this goes. At $7k a year or more per eng in token use, I think it's very reasonable to buy engineers machines with obscene GPUs and RAM and run models locally. And if it doesn't make sense now, someone will figure it out and save companies $10k+/eng over 3 years.
- charcircuit 5mo agoThat could leave idle time where GPUs are sitting unused. It would be better to have a shared cluster that many engineers all share. And to avoid a cluster not being saturated other companies queries could also be batched. And oh wait we are back to doing AI inference in the cloud as it is an efficient way to serve AI.
- no-name-here 5mo agoIf you only want/need the kind of model output that can be served on a machine costing single digit thousands, aren’t cheaper cloud-served models available? (And as the sister comment points out, sharing hardware allows greater utilization and lower costs per user.)
- Glyptodon 5mo agoThat might just mean even more savings as you'd only need need a size n cluster for m engineers where n is probably < m.
- slopinthebag 5mo agoI imagine there are companies forming now with their entire business model being building "prosumer" inference machines and farms running everything from Qwen 3.6 27b up to GLM 5.1 and everything in between, packaged perfectly for companies to make one-time investments in with the assumption that open models will be getting both more efficient and better over time.
- bananamogul 5mo agoThe good news is that this might be the end of Oracle.
- JohnMakin 5mo agoI've sort of lost some respect for ed that I had early on in the hype cycle - he's still right about some things, but I can see him slowly and subtly retreating from his strong position, held even a few months ago, that these things will never ever be useful for anything and it's all a scam because they don't actually do anything at all except burn money. He would say it like 8 times a monologue. I remember one podcast maybe ~6 months ago he brought a developer skeptic on, and was trying to get him to say it wasn't actually useful for coding, and the dev was like "maybe not as advertised, but I definitely use it and it is useful to me" and he pivoted off the topic very quickly. It seems he realizes he was wrong about that and has pivoted slowly to, "well, maybe they work sometimes, but the cost isn't justified." Which is a reasonable question! I just find his style of never admitting when he is wrong off putting and the way he presents things as absolute fact, when he's guessing like the rest of us. He was right about a lot, wrong about a lot, it's okay to admit that, I don't think his fan base would care.
- hparadiz 5mo agoThe economics is spending a few hundred bucks on software for an IC you're already paying over ten grand a month in order to make them more productive. How are supposedly smart industry experts not seeing this obvious fact? Are these guys actually experts?
- xienze 5mo ago> The economics is spending a few hundred bucks on software for an IC you're already paying over ten grand a month Let's be fair here, the endgame is not "a few hundred bucks a month." Not for how much money has been invested. How much extra you have to spend to make developers how much more productive, and will companies go along with it is the trillion dollar question.
- hparadiz 5mo agoYou know I can just lookup the costs per seat right? It's not that much and not everyone is a heavy user at an org. And for code the costs are falling per compute cycle.
- pmdr 5mo agoI wonder how long until this post is flagged/"shadowbanned". Such was the fate of almost all of Ed's posts on HN, with little surprise as to why.
- CamperBob2 5mo agoPeople who don't adjust their prior outlook in light of newer data may not be the best fit around here. I'm OK with that.
- pmdr 5mo agoWhat is the newer data?
- margalabargala 5mo agoExtensively discussed elsewhere in this thread. Just start at the top and start reading comments.
- maplethorpe 5mo agoCan you summarise? I only reached your comment after scrolling past all the others and I still don't have the answer. Is the new data that models are more useful for coding than they once were?
- margalabargala 5mo agoThat sounds like a reading comprehension skill issue? In which case I don't see why me summarizing would move the needle. But if it helps, no, the data being discussed is surrounding the economics of running inference and R&D, nothing to do with the utility of models for coding.
- maplethorpe 5mo agoYours is the first from the top to mention this. You might want to consider the physical location of your comment before telling people to read the thread. We could do without the rudeness, too.
- OrvalWintermute 5mo agoI think the company Taalas alone destroys Ed’s arguments Because, comparing vs GPUs ~16k–17k tokens/second per user <1ms latency 10x power efficiency 20x cheaper production Model to Si ~ 60 to 90 days We have every reason to believe SW_to_Si will facilitate improving economics
- BosunoB 5mo agoAll subscription models are subsidized by users who don't use much. The fact that somebody on a $20 sub might get $50 in value isn't crazy if there are 3 people who only get $10 in value. This isn't some sign that the model is broken, it's the intended outcome. Also, I didn't read this whole thing, but I have yet to see Zitron respond to the strongest AI financials claim, which is that the models themselves are profitable on a life-cycle basis, even if the companies are not profitable on an annual basis due to capital expenditure. Dario made this claim exactly, and it more or less blows all of Zitron's financials arguments up.
- CodingJeebus 5mo ago> which is that the models themselves are profitable on a life-cycle basis, even if the companies are not profitable on an annual basis due to capital expenditure. Until they file an S1 to go public and show the world the books, take everything they say with a grain of salt. The amount of financial engineering going on in this space is astounding, and I'll believe it when I see an objective 3rd party release an audit confirming this claim.
- weakfish 5mo ago> but I have yet to see Zitron respond to the strongest AI financials claim He does in this [0] article. [0] https://www.wheresyoured.at/ai-is-really-weird/ https://www.wheresyoured.at/ai-is-really-weird/
- BosunoB 5mo agoThanks for the link. I'll admit I'm not an expert on the business side of this, but is this really much of a response? He seems to just call it strange accounting and then he moves on. It doesn't even feel like particularly strange accounting to me. Aren't there plenty of companies that spend a lot in one year and realize the gains in the next year? If I build a house this year and sell it next year, the house was still profitable, even if next year I'm building 3 more houses to sell in the year after.
- csande17 5mo agoZitron has responded to that claim here: https://www.wheresyoured.at/ai-is-really-weird/#does-anthropic-measure-its-gross-margins-based-on-how-much-revenue-a-model-made-rather-than-revenue-minus-cogs https://www.wheresyoured.at/ai-is-really-weird/#does-anthrop... The TL;DR is that Dario likes to talk about imaginary/hypothetical companies a lot in interviews, and those companies' financials don't have a direct basis in reality.
- chankstein38 5mo agoBefore subscribing to Claude, I put $15 into my account so I could use it from Cline in VS Code. After less than a few hours I was out of money. This was basically just to get a simple project setup and a few 1000~ line (AI generated) code files edited. I have heard Cline is less ideal with token management but regardless, these services can easily cost us hundreds or thousands of dollars a month billed on usage. ($15x4hoursx2 for a work day = $30, $30x25 = $750). And that is assuming my very light usage here could even apply to a larger code base. My guess would be if I hooked it up to an enterprise project it'd skyrocket easily to $60+/day.
- ludicrousdispla 5mo agoDoes this mean we can just go back to using software libraries?
- gwbas1c 5mo agoWhat's the quote? > Don't attribute to malice what can be attributed to incompetence. We're currently used to SAAS billing models that are either all-you-can-eat subscriptions, or metered around some easy-to-understand metric like # of users, or otherwise number of gigabytes consumed. The SAAS economics work that way because the compute consumed is typically too cheap to meter. Some customer uses a little more than average, some customer uses a little less than average; it's not worth the time to even it out to the penny. AI is so darn CPU (GPU? AIPU?) intense that will only be profitable, and affordable, if it can be metered like electricity and billed with a small margin. In SAAS, we're not used to metering billing computations this way.
- mitjam 5mo agoI would be curious to see a calculation backwards from TAM. Napkin: 50M developers worldwide (SlashData, 20M in China and India). If every developer had a $200/month subscription, that‘s $10B / Month. I think, many developers are expected to pay much more than that.
- Lionga 5mo agoMost developers in China or India have a monthly salary of 1 K USD. If you expect them to pay way more then 200USD thats like asking US Devs to pay 5K a month. Yeha not gonna happen. And the funny thing is the estimate pure CAPEX Spend of AI companies needs them to earn about $20B to $40B a month to cover cost of capital alone of their trillion dollars of investments.
- warkdarrior 5mo ago> Most developers in China or India have a monthly salary of 1 K USD. If you expect them to pay way more then 200USD thats like asking US Devs to pay 5K a month. Yeha not gonna happen. That's exactly what is going to happen. India/China prices will be $100-200/month, US prices will be $5000/month. Keep in mind that most of these costs will be covered by the employer. It'll put downward pressure on dev pay, of course.
- warkdarrior 5mo agoMicrosoft made $17B/month in 2024, Google made $25B/month, Amazon $48B/month. And the computing market is growing.
- aaroninsf 5mo agoEd, my friend, I've got some news for you. Economics Don't Make Sense. I mean, seriously... our current late-stage capitalist economy is the chaotic sloshing of excess capital or inverted debt in a shallow tub within which clumsy giants are stamping like toddlers, and a parasitic kleptocratic oligarch class balances its efforts biting the toddler ankles in hope of more stamping judged advantageous, and, bagging what water they can.
- jsLavaGoat 5mo agoEd could have been right, but I think he's a bit of a front runner than ended up being out too far and not accepting that, for coding at least, the tool is useful. And coding is a big business itself. Of course there are always going to be shenanigans to point out, and I'm glad there are skeptics.
- fancyfredbot 5mo agoHe does have a point about fees. It's not really surprising that the fee structure designed for chatbots would not make sense when applied to long running tasks and agents. But an increase in prices can solve this problem. Doubtless some people will reduce usage as a result. But Ed seems to find the idea that a 10 man developer team might spend 80K a year on tokens ridiculous. I don't understand this. Has he seen how much developers are paid? If you get a 20% productivity boost from coding agents, then that's two developers for 80K - effectively very good value. Where things could go wrong is in comparison to cheaper models. If it's 5K a year for Qwen, and it's 2/3 as good will you pay 75K extra for Opus? Perhaps not.
- blks 5mo agoI think that team is better off with a junior developer. This alleged “20% productivity boost” even if it exists, is individual. On the team level, it will be largely offset by people having to review 20% more code.
- fancyfredbot 5mo agoObviously in some cases a junior developer is a better investment if it's a straight up choice. Actually I think it'll be rare for a manager to be choosing between either a junior developer or a coding assistant, since each are going to benefit the team in very different ways and it'll often be obvious which you need. What I mean is that at the price levels in the article the coding agent still had a realistic chance of positive ROI. People will pay for things with positive ROI.
- Yizahi 5mo agoThe problem is that LLM cost is more or less the same for generating some fixed amount of code or it will converge to that soon. But developer costs vary wildly based on the seniority*geographical location. Sure some Silicon Valley architect will be always more expensive than any LLM bills he incurred. But a middle tier dev at an outsource or local cheap shop overseas using the same LLM for the same tasks and same token costs? Eeh, it can go either way really.
- thinkindie 5mo agoi don't know when it was introduced, but Claude Code has recently added the cost for your session when you run /usage
- gnachman 5mo agoSo you’re saying there’s a chance that Oracle will die? Sign me up.
- purplepatrick 5mo agoI keep seeing articles like this that extrapolate from token pricing onto token costs. This is wrong. Companies don’t sell their goods/services at cost. A model’s being priced at, say, $30/M for output tokens doesn’t say anything about what it costs the company to provision the 1M tokens via the model. And no, you cannot extrapolate from any company margins that someone may have overheard in an SF coffee shop onto individual product line margins or their trajectory either. This information is usually unknowable even in most SEC filings for public companies. It’d be great if people who wrote these articles used, say, AI to look up some basics on how a business operates. It’s really easy to do, believe me ;)
- alok-g 5mo agoA possibly naive question: Mobile phone plans, Internet service providers, etc., also often used fixed monthly pricing. How does that keep working (with competition present)? Is the issue monthly pricing, subsidies, or both?
- latentframe 5mo agoThis seems a classic capital cycle problem, with huge upfront investment, unclear pricing power and everyone scaling supply at once and so that combination usually doesn’t ends with great returns
- kk_mors 5mo ago[flagged]