4 ms·
There's a few major problems with the article. The most obvious is that frontier labs are not charging remotely close to the cost of tokens; afaik most estimate
by joshjob42 5mo ago
There's a few major problems with the article. The most obvious is that frontier labs are not charging remotely close to the cost of tokens; afaik most estimate north of 80% profit margins. As a reference, providers are profitably providing Kimi K2.6 for $4/1Mtok out. Is that as good as Opus? No, but it's probably at least Sonnet level, so that's ~4x cheaper than Sonnet while still being profitable to serve on the margin. So you aren't plausibly getting into actual subsidization territory until you're over 5:1 sub to nameplate token costs.
How many tokens can you realistically burn through in one chat session? Opus and many other frontier models do maybe 60tok/s, less 250k/hr out. In you can use more, but in most cases cache is 5-10:1 cheaper than new input. Say you average 500ktok in, 90% cache, per request. That amounts to 100-150ktok in new input-equivalent costs, which in most cases is ~20-30ktok in output-equivalent costs. Do a request every minute, that's a total of about 1.5-2Mtok/hr. At API prices that's $50/hr for Opus, but really it probably only costs Anthropic $10/hr to serve that.
That said, even if a developer is burning $50/hr, many, many employees at large companies cost more than $100k/yr to employ all costs considered, so making them say 20-30% more productive can easily make that worth it for most. If the labs shave their margins ultimately to more like 20-30%, you'd have ~$15/hr in costs to use the services, and nearly every white collar job is way over 30k/yr to employ. If your salary is 80k, you probably cost the company 200k all in, so making you 15% more productive offsets the $15/hr cost.
So first party providers are not in a horrifying position or anything from a subsidization standpoint. The people in bad shape are Cursor and Perplexity, who don't have frontier models and are dependent on the open source community, which is typicly 6-12 months behind the frontier. They have to pay full freight API costs at 80% margin for the big boys to serve their harnesses, which is indeed untenable, and they'll have to either force users to use open source models and/or in house models they can serve at-cost or they will have to charge vastly more.
Gemini, Claude, and ChatGPT first-party services like Antigravity, Codex, and Claude Code are not in serious trouble though.
- ToucanLoucan 5mo ago> That said, even if a developer is burning $50/hr, many, many employees at large companies cost more than $100k/yr to employ all costs considered, so making them say 20-30% more productive can easily make that worth it for most. If the labs shave their margins ultimately to more like 20-30%, you'd have ~$15/hr in costs to use the services, and nearly every white collar job is way over 30k/yr to employ. If your salary is 80k, you probably cost the company 200k all in, so making you 15% more productive offsets the $15/hr cost. Nobody including the connected article is making the argument that this cannot be profitable ever. People are saying "there is no way this admittedly quite interesting tool is going to be able to make back all of this money" and I think they are completely right to say that. You can absolutely make money with this stuff, just not at this scale. The buildout for this shit has been certifiably crazy and a number of the involved firms are overleveraged for tens and even hundreds of billions of dollars. How in the sweet fuck are you paying that off, plus giving investors dividends, selling this at $15/hour/user??? That math does not math. A quick google says there are between 1.5 and 4.4 million developers in the US alone, let's say it's 5 million, to be generous, and each of them is subbed to this for 8 hours per day, continuously. That's 600 million per year in revenue. If you took ALL that revenue, and put it towards paying down this debt, not leaving any for employee salaries, upkeep, ongoing development, it would take DECADES to pay down what OpenAI already owes. And yes I'm sticking directly to code, because that's the only thing I've seen it be really good at. Are we really proposing that every knowledge worker on earth and every manager of such workers is going to have an autonomous agent running all the time!? To do what, make sure they don't have to read or write email? Which even just that example is bringing in a fucking mess of legal, compliance, and security violations because LLMs are not intelligent and are not capable of being properly secured. Like I'm sorry, I cannot take this industry seriously when even the most basic back-of-napkin math is saying, nay, screaming from the rooftops that they are FUCKED.
- vidarh 5mo agoBy your numbers, it'd be $120/day per developer * 5 million = $600m per day, not per year. Of course people don't work every day, but even with European-level holidays that number is off by a factor of 240 or so.
- ToucanLoucan 5mo agoQuite right, honestly not sure how I fucked that up so bad but I'll own it. Okay so all we need is every coder + 0.6 million more or so in the United States, subscribed to this for 8 hours a day, and the business model can work. That still feels incredibly optimistic given how split the community at large seems to be about how good this tech is, and it assumes all those developers also all work for firms large enough to pay for all of that. However we are still very much in back of napkin math. We haven't even gone into what it costs to provide these services, how much it's going to cost yet for all these datacenters to be built, how much electricity and water they're going to rip through, their own employees and basic overhead, and all the rest. So IMO, we've now elevated it from "hopeless" to "this could work if a whole lot of other things line up really well."
- deleted 5mo ago[deleted]
- asdfasgasdgasdg 5mo agoIt's not just developers who are using this. My economist friends are. I bet most business analysts and general administration folks are or will be soon. Every normal person I know in my neighborhood is using AI for this thing or that. 50M people are currently subscribed to ChatGPT and it would be very surprising if this number goes down in the future. I dunno I think about the language some people are using about AI investment and it is reminiscent of the many years where people were saying Amazon was a bad buy because they never turned a profit. Admittedly AI companies are investing more than the money they've already brought in, but I would be very hesitant to predict that it's all froth given the usefulness I've gleaned from the tools. Don't get me wrong, I'm not unconcerned, but I think there are good reasons to suspect that at least some of the AI companies are making sound investments.
- loeg 5mo ago> How many tokens can you realistically burn through in one chat session? I've used single digit billions in a couple days, FWIW.
- bwestergard 5mo agoWhat sort of work were you doing?
- loeg 5mo agoConverting a couple hundred kLOC C++ codebase to Rust.
- bwestergard 5mo agoCool. Sounds like it went well?
- loeg 5mo agoMaybe! Still evaluating if the output does what it's supposed to do.
- xienze 5mo agoNot the parent, but the way developers are basically trying to create entire development "teams" consisting of multiple agents that work around the clock using the latest, most expensive models (naturally) lends itself to burning insane amounts of tokens.
- kcartlidge 5mo agoI'm a fair bit lower than some others as I only use it outside of work hours on my own small projects, but my Cursor account shows (for a random recent date) 12,184,233 tokens in a day. That day feels pretty representative. That's with 86 interactions spread intermittently over a couple of hours so if I did a full working day like that I'd be looking at maybe 40 to 50 million.
- doctorpangloss 5mo agolots of words. do you think per token prices will go up or down in the long term? will the price per task trend down or up? what about the price of human labor?
- GardenLetter27 5mo agoThe price of everything will go down. That is the beauty of the free market.
- rspeele 5mo agoIf the price of everything would go down it wouldn't be too concerning and everybody would be on board with the "beauty" of it. What seems to actually be happening for white collar workers is that the price they can charge for their labor is dropping, but the price of their expenses (housing, food, gas) continues to rise.
- dgellow 5mo agoThe free market hypothesis is about resource allocation, nothing to do with price of everything going down
- Yizahi 5mo agoIn the absolutely free market price will go up a lot in the end. Because only one monopoly will exist by that time and it will jack up prices to the maximum tolerable level. And that level can be surprisingly high, because in every human activity there will be few willing to spend crazy amounts of money for practically anything they perceive valuable.
- mike_hearn 5mo agoThis kind of argument relies on odd definitions of "truly free" that boil down to anarchism, which isn't what anyone who advocates for a free market means.
- 5mo ago
- zozbot234 5mo agoIt's not even a fixed cost per token (even though it's billed that way, and that's still miles better than a fixed-price all you can eat). You're incurring a cost that's proportional to generated tokens times the context for each (plus the prefill cost for any uncached input), so the expense grows quadratically with your average generated context. This all becomes extremely visible when trying to do agentic coding with local language models - you quickly realize that controlling context length and model size is just as important as avoiding wasted effort. The real scam is not AI Q&A ala ChatGPT, that's actually quite viable - though marginally less so as conversations grow longer. It's agentic coding with SOTA models and huge contexts.
- GaggiX 5mo agoUsing larger contexts often costs more in the APIs or consume more of your quota but this is becoming less of a problem with models using more clever attention mechanisms and not just full attention on all layers. You can look at: https://sebastianraschka.com/llm-architecture-gallery/ https://sebastianraschka.com/llm-architecture-gallery/ and see how much things have changed.
- margalabargala 5mo agoThis is also something of a non issue because as context grows and attention gets diluted, the models perform worse. It'll cost Anthropic more to run your 900k context session, yes, but it's in your interest not to have a 900k session in the first place.
- great_psy 5mo agoYou’re right about performance degradation, but good luck trying to sell that as a product. You can drive this car, but the last mile of this trip will use as much gas as the first 20 miles. I think it’s in anthropics interest to keep this fact hidden from CEOs who push for ai adoption.
- intended 5mo ago> afaik most estimate north of 80% profit margins This seems to be the lynchpin of your argument. It makes me wonder if I have been living under a rock, because I have never heard of frontier labs making money. AFAIK all AI firms are simply burning money to acquire customers at this stage. Is this wrong?
- asdfasgasdgasdg 5mo ago>It makes me wonder if I have been living under a rock, because I have never heard of frontier labs making money. You're confusing the profit from the marginal token and overall profit (basically gross margin and operating margin). The comment you're replying to is calculating that AI labs are probably making a substantial profit per paid token. It's just that so far that profit has not been able to overcome the ongoing R&D and capex costs.
- kgwgk 5mo ago> not been able to overcome the ongoing R&D and capex costs. And the cost of not-quite-paid tokens.
- margalabargala 5mo agoWhich may or may not exist, hence this thread.
- kgwgk 5mo agoNon-paid tokens do definitely exist and they weren’t included in the remark about “substantial profit per paid token”. Underpaid/subsidized tokens also exist which don’t provide “substantial profit”.
- margalabargala 5mo agoAre you talking about free promo tokens the company gives out, or are you implying that subscription tokens are sufficiently subsidized so as to be below cost?
- boelboel 5mo agoIsn't this akin to saying Big Pharma companies could easily make money if they just stopped doing expensive research? The massive R&D spend is the core of the business plan; it's the only reason they can demand high prices in the first place. Once OpenAI stops spending billions on training, their pricing power vanishes because users will just migrate to Anthropic or whoever releases the next frontier model. Would imply there'd be space for only one to outlast them all in some sort of war of attrition (perhaps similar to silicon industry).
- kimetime 5mo agoBig Pharma does seem like a good comparison for frontier lab business model. Doesnt really have the patent protection or distinct diseases pharma does, Wonder if labs start more heavily branding “specialties” instead of general capabilities to develop some differentiation
- tobbe2064 5mo agoYour math is pretty bad 50$/h is a yearly cost of going by swedish standards, 50$/h × 40h/week × 48 weeks / year = 96k$/year At that rate is a really shitty bargin for 30% increase in productivity. Even if you drop it to 20$/h and sort of break even, you are loosing competens building and teory building, decreasing the likeleyhood of making architectual progress and risk getting bogged down in a swamp.
- joshjob42 5mo agoAn employee often costs a company 2-3x their salary, so someone making 100k a year, costing 300k/yr, who is made 33% more productive (100k more worth of work to the company) offsets the compute cost.
- lbreakjai 5mo agoProblem with this math is it always assumes some ridiculous baseline compensation (or costs, in this case) as a matter of fact. There's an entire world of developers not costing 200k to their employers. Truth of the matter in most companies large enough is if you make your devs 30% more productive, then that'd mean 30% more code going through "change management" hell for months. You're not even paying to stand still, you're just pushing even more down a bottleneck. The price most people are willing to pay to make things worse is close to zero.
- Balinares 5mo ago> providers are profitably providing Kimi K2.6 for $4/1Mtok out. Do you perchance have a source for this? Is the profitability assessment comprehensive, including hardware amortization? I've found it hard to track down actual hard numbers for the cost of inference.