5 ms·
Price per 1M tokens is meaningless
- zeroonetwothree 3mo agoWell, not totally meaningless but certainly can be misleading.
- deleted 3mo ago[deleted]
- tidbeck 3mo agoRelated to this, for our use case, setting thinking to high instead of low made tasks complete faster and cheaper (Gemini 3.0 flash). Other aspects are caching, often at 0.1X cost, where providers really differ in how efficient they are (Anthropic really good, Google not so much) and how chatty a model is (costing output tokens).
- janilowski 3mo agoI also don't see much of a quality difference switching between thinking levels while using those models as agents.
- shay_ker 3mo agocost per benchmark task is definitely interesting! i've always wanted cost per prompt, but even that has too much variation.
- janilowski 3mo agoWell yeah this benchmark I've quoted is quite large, I believe they use thousands of diverse tasks, so this average cost per task could be a more accurate representation of how expensive each model actually is to run.
- vfalbor 3mo agoI believe the future lies somewhere in between. I'm working on a hybrid application to reduce our company's token consumption. It runs on our data center's computing infrastructure and on laptops in our community. You might be interested; you can check out the code if you're interested: https://github.com/vfalbor/hibrid https://github.com/vfalbor/hibrid
- BugsJustFindMe 3mo agoPrice per token is meaningless for more reasons than this, because all of the provider monthly subscriptions price tokens _extremely_ differently than their per-token billing rates. It's stupid to look only at what you get when paying more than you need to for a given service.
- janilowski 3mo agoYes that's true, although calculating the exact discount you get using a subscription could be difficult, especially since the labs often don't state their actual limits and just tell you you get "more usage" or "even more usage" or "5x the usage" (but how much exactly is x?). It is clear we are getting a very meaningful discount by using a subscription though. I once checked how much would I pay for Codex after about 3 months of non-daily use if I only bought "credits". Couldn't believe it when it came out to be north of $1000 (I paid them $20/month in that period).
- BugsJustFindMe 3mo ago> it came out to be north of $1000 (I paid them $20/month in that period) Yeah. My $200/month subscription would easily cost $5000/month or more if I were paying per token.
- dandaka 3mo ago> calculating the exact discount you get using a subscription could be difficult Why can't you derive this discount from experiments? Example of such research https://she-llac.com/claude-limits https://she-llac.com/claude-limits
- yreg 3mo agoCost per token doesn't say a lot, but "Cost per benchmark task" is also meaningless if your task is difficult enough that the cheaper model has no chance of cracking it.
- janalsncm 3mo agoSimilarly, tasks that are too easy also aren’t ideal either. If a small model makes mistakes and backtracks but eventually cracks it, it will be using a lot more tokens than a bigger model that does it all with minimal mistakes.
- sweetjuly 3mo agoI think what you're really getting at is that it's only useful if the benchmarks are predictive of your workloads. If it predicts well (for example, your tasks are equally easy), then the fact that a larger model can complete it more quickly means that you may be able to complete the task more cheaply, depending on the token cost. If the benchmarks are non-predictive, well, you can't use them for much of anything, which is of course a recurring problem with every benchmark ever.
- yreg 3mo agoYeah, if the benchmark is actually predictive of the tasks you have then it is trivial to conclude that the cheapest-per-benchmark-task model will be the cheapest one for your tasks…
- janalsncm 3mo agoIt might vary between tasks though. A model that’s great at abstract reasoning might be great at writing math proofs but struggle to write software in <insert language>.
- efromvt 3mo agoIsn't the benchmark working exactly how it should in that case?
- lifeisstillgood 3mo agoMy advice to any CEO / individual - throw your hands in the air and bring it in-house. Yeah the performance can dip depending on what GPUs you can salvage these days but the uncertainty over price is almost nothing compared to the uncertainty over the effective use of AI. It’s not just coding (do I go partly agentic or all out Steve Yegge). This is all over the enterprise - do we parse every email, rewrite PowerPoints? Or just stop using PowerPoints at all. Do we throw LLMs at the mess of wikis and word docs, do we pretend that the policies no-one has ever read actually are how the LLM should think or is it how the work actually gets done - barely documented The uncertainty of how to use this vastly vastly outweighs the price in a data centre - so buckle up, buy enoughbGPUs to experiment at a known cost and one day you will find the approach that gives you 10x returns - at that point pay any price per token but not till then
- nojito 3mo ago>bring it in-house People don't like to hear this but the open models just aren't good for end to end agentic workflows. There are some very very good small open models that can excel in certain finite bounded tasks, but the foundational models are essential to building out agentic pipelines that actually work.
- TacticalCoder 3mo ago> People don't like to hear this but the open models just aren't good. Stuff like the latest DeepSeek, Kimchi and GLM are used and loved by many people. It's not using an open model that is difficult: it's having the hardware allowing to do so. It's pricey and require technical skills. That's why most people who are using (excellent btw) open-weight models are just renting compute online.
- nojito 3mo agoThey just aren't good at agentic work. Also risking it all for some distilled models is a recipe for disaster.
- koolba 3mo agoThis reminds me of cpu benchmarks vs actually running games and measuring FPS.
- dbuxton 3mo agoAs well as cost-per-task I think it's worth thinking about speed, especially in non-coding contexts that benchmark less cleanly We've started trying to do some comparison videos to capture more of the UX vs speed vs cost stuff e.g. https://www.linkedin.com/feed/update/urn:li:activity:7479891298603282433/ https://www.linkedin.com/feed/update/urn:li:activity:7479891... which one of my team did for my LinkedIn account (disclaimer: marketing) (In this particular case Deepseek was way slower than GPT 5.5 but I think that's because it installed Libreoffice half-way through the task!)
- sillysaurusx 3mo agoConcrete example: I’ve been trying to use Claude to generate all my commit messages, but it takes 5-10x longer than if I just write them myself. Mine are less detailed, but one line changes are sometimes inconsequential (especially white space reformatting). I wish there was a model that understood the codebase well enough to generate commit messages in half the time.
- janalsncm 3mo agoPricing based on tokens always seemed a little weird to me.“Tokens” was and still is an engineering concept. The fundamental unit of transformer encoding and decoding. But I have a sinking feeling that many AI developers think “tokens” got their name from the same idea as “virtual tokens in a casino” which is more related to product pricing and business.
- pooploop64 3mo agoTokens in a casino is pretty accurate if you think about it. You never really know what you'll get so it's tempting to "roll" over and over, thinking every roll puts you closer to a bellringer. It can even get addicting for some people.
- HarHarVeryFunny 3mo agoTokens do reflect the provider's cost though - each token output required them to execute the model once, normally incurring a fixed amount of compute per token.
- janalsncm 3mo agoProviders amortize the compute across a batch. If yours is the only request in the batch it will cost them one full pass through the model. If yours is one of 1024 inputs in the batch the per token cost is 1024x less.
- HarHarVeryFunny 3mo agoApparently hardware depreciation is the dominant cost rather that operational cost (electricity etc), and this is occuring at a fixed rate per the planned replacement lifetime. So, the cost to provide the service is essentially fixed regardless of load, but the revenue they are generating is variable. In practice most GPU's are going to be capacity-maxxed since the providers sell cheap batch APIs that they queue to keep the machine loaded. They'd be losing money over a given time interval if revenue generated during that interval wasn't greater than depreciation (etc), but it seems that will rarely happen.
- nathanyz 3mo agoThe Sonnet 5 comment is spot on. Even Anthropic's own graph initially showed lower performance at higher costs. Only thing I notice about Sonnet 5 is that it does appear to hand off tasks to agents more frequently similar to Fable, but of course nowhere near the quality of Fable. My guess is that Opus 5 will do similar but just isn't ready yet.
- cute_boi 3mo agoSonnet 5 is a huge regression and many times it performs worst than deepseek. I believe Antrophic staff themself don't use Sonnet and use Fable for everything.
- asdewqqwer 3mo agoGiven the capability of fable and the shockingly repetitive silly mistakes they made when publishing/updating something, I am starting to wondering whether Anthrophic can afford Fable for everything themselves.
- teravor 3mo agotool use is another factor, every time the agent uses a tool the entire context is priced at cache rate on top. the same happens when it asks you for input.
- Jcampuzano2 3mo agoI keep trying to convince directors and executives at my company to look past the cost per token amount but they refuse to do so. Those are the only things that actually give any sort of measurement of the monetary value of a token by these labs, and so its what many go by. For example there's some benchmarks that show that Opus for any task that requires a higher than `high` level of effort, may have actually been cheaper to use Fable on low even though the cost per token is drastically higher Similarly with GPT 5.5 vs Opus. They simply look at the dollar amounts the labs assign to each model and run with it. But part of the issue compounds on the fact that there are many people who simply default to the smartest model/effort and don't actually vary their model per task. So in some sense I don't actually blame them very much.
- greenavocado 3mo agoThat's why you have to reframe in terms of total cost per task and factor in model token generation quantity and multiply that by the base cost of the model. Then factor in your time value if you dare. Then you should get a more meaningful business metric.
- xienze 3mo ago> may have actually been cheaper to use Fable on low even though the cost per token is drastically higher Well that's the problem with these black boxes. You really have no idea beforehand how many tokens a given task is going to take. There's simply too many variables involved. It's therefore only natural for people to assume "the cheaper and older model is probably going to cost less overall to use than the newer, more expensive one."
- Schiendelman 3mo agoI'm definitely using more credit per task, using Fable, but I'm also giving it much more difficult tasks. You're right, it's very hard to tell. I think that using Opus would be more expensive because of the number of iterations I would need to go through.
- Schiendelman 3mo ago
- kpw94 3mo agoIn the context of local LLMs on limited hardware I've ran to the exact same conclusion: "tok/s" isn't the most useful metric when my personal North star metric, given my fixed hardware is: Model smart enough to execute my goals _in the minimum amount of time_. Some models I tried (Mistral I think) had better tok/s, and roughly same billion parameters / scores on various benchmark... But they were _so_ verbose, that they generated many more tokens compared to a Qwen model of same caliber to answer the same thing. So even though it had better generated tok/s, because so many more were generated, the clock time was longer. And this compounds over mutli-turns: more generated token means more context used in the next turn (until some compaction or something runs)
- c7b 3mo agoEven more important in a local context is the difference between token generation and prompt processing speed. We tend to focus on the former, but for multi-turn/agentic workflows the latter can dominate.
- kpw94 3mo agoYeah definitely. I've recently commented on that: https://news.ycombinator.com/item?id=48557890 https://news.ycombinator.com/item?id=48557890
- PunchyHamster 3mo agoI feel like we need to see more proliferation of local LLMs to start seeing ones turned to be terse, rather than maxing the amount of tokens user pays for
- sleepybrett 3mo agowould be nice to have these benchmarks so they can be run against models like the qwen family, gemini, etc.
- Cappybara12 3mo ago[dead]
- shireboy 3mo agoYup. I’ve been evaluating several on openrouter and find token cost meaningless for my work. I haven’t found a great alternative, though the “cost per task” he uses makes some sense.
- danielmarkbruce 3mo agoAn LLM is an extremely complex thing used for all manner of purposes. The hope that there would be some simple pricing construct that would map nicely to value provided is a pipe dream. Pricing per token is at least reasonably straight forward. If you aren't getting value, you don't use the service. One doesn't buy a Ferrari and then complain that in their town Ferrari doesn't help them pick up women and hence it should cost less.
- janilowski 3mo agoI'd say it's more like going to a Ferrari dealership and they tell you they will build a car for you and bill you per gram of parts used. It might also not work. And they might also not build it ever, really — but they will bill you for any attempts to build it.
- danielmarkbruce 3mo agoThat's fine though. If that doesn't work for you, don't buy. There are all manner of situations where what one wants or needs and what they get don't match up well. You don't price out every situation - it's take it or leave it. Pricing in a way that is somehow based on cost structure at least enables the provider to work to reduce the cost and hence price and win. Costco prices at a small margin above cost, they don't price "if this meets value prop X, pay Y and if only value prop A, pay B".
- dandaka 3mo agoAnother important benchmark would be — cost per benchmark task using subscription tokens. Since most of us are using subscriptions and cost per token there is quite different from API costs.
- janilowski 3mo agohttps://news.ycombinator.com/item?id=48821204 https://news.ycombinator.com/item?id=48821204
- Lerc 3mo agoCost per tokens is as valid as price per unit volume of fuel. Changing the fuel type, efficiency of your vehicle, driving distance, or driving conditions will all change how much it will cost you. Fuel cost per unit volume does not become meaningless just because you are neglecting all of the other factors involved. That would be throwing away the only data point you have been using. This is just asking for someone to amalgamate all of the factors involved into one simple, easy to game, index.
- bjt 3mo agoExcept, a gallon is a gallon no matter which gas station I'm at. Also I know my car's gas mileage, and it doesn't change when I visit a Shell station instead of a Chevron. The composition of the gas is regulated, as are the pumps that dispense it. There are inspectors from the state whose job it is to ensure that when I buy a gallon, I really get a gallon. Tokenizers aren't standardized to anywhere near that level. A "token" from one isn't the same as a token from another.
- sieabahlpark 3mo ago[dead]
- Aurornis 3mo agoThat’s not a good analogy because a gallon of gasoline has a known amount of energy in it. The efficiency of each vehicle is also known, at least in a way that is easy to compare on a relative basis. I can go to 10 different gas stations and buy the same amount of energy from them. When I put it in my car I’m going to get the same result out. The differences are very small.
- mikebs1 3mo agoThe variable missing from cost-per-task: which tasks shouldn't be hitting an external API at all.
- kfsone 3mo agoI feel we are caught in a "this is fine, pay more and we may turn down the fire" situation. The LLM itself produces one token. Some tool adds that token to the input and runs it again, flogging the horse. Downstream another tool, some kind of harness, tries to control this stream by injecting tokens into the context and then sending it to the inference tool, and then trying to pattern-match the output. Finally, there you are on CodePorn.yata paying for an agent to generate code, paying for an agent to tell you what's wrong with it, and paying for an agent to make it differently bad, and hopefully move on to the next task. If it still hasn't dawned on you that this isn't just a bubble, but a snake-oil-bubble-bath, just try to imagine the paradigm shift whereby you go on github.com, assign an issue to an agent, the agent fixes it by rewriting the application in Pascal but a reviewing agent catches that you wanted it to print a measurement in Pascals (pa), and you don't pay for the work or the review, you only pay for work that one or two reviewing agents determine is up to par. Nobody is going to do that because as soon as they test it they're going to have to do some math that won't make sense without admitting/realizing it's not some near-sentient, AGI rating 0.9 intelligence, it's just a text prediction algorithm that can pull out entire sentences when you use it to infer output on topics it trained on.
- c-hendricks 3mo agoOne more ~~lane~~ layer of LLMs is sure to solve all our problems
- paulddraper 3mo ago> it's just a text prediction algorithm that can pull out entire sentences when you use it to infer output on topics it trained on This downplays the incredible things that can be done with it. There's a lot of noise, yes. How long has the web existed? And yet we're still figuring out how to optimize (HTTP/3). Disregard the signal at your own expense.
- EA-3167 3mo agoWhat incredible things can be done with it?
- dools 3mo agoIt’s not meaningless at all: every query returns usage and I can calculate the cost. EDIT: this is like saying hourly rate or salary is meaningless. Different people have different output. You have to evaluate performance. EDIT2: just pray the LLM providers don’t start taking Patrick McKenzie’s advice and start charging based on “value delivered”
- bArray 3mo agoThe only metric that really matters is 'profit per amount invested'. This is very difficult to quickly evaluate, and therefore we resolve to use simplified metrics such as cost per unit tokens. The point at which the metrics become meaningless is when others become aware of them, and begin to optimise for them. Lines per code is is not a bad insight for development activity, only when the developers are not aware of the metric. Price per 1M tokens became meaningless when LLM providers started to optimise for it. It seems to be that Sonnet 5 is optimised to score well on AA intelligence whilst seemingly having a lower price per 1M tokens. I think generally we are in an AI bubble, and it will at some point pop. The numbers simply don't make sense. I would gamble heavily on local cost per task to survive the LLM winter. Given that hardware is pretty much a fixed overhead, you probably want to optimise for task per kW - that's where I'm betting.
- deleted 3mo ago[deleted]
- usef- 3mo agoSonnet 5 makes more sense when you pretend the higher thinking efforts don't exist. (His test was on xhigh) Anthropic's own release announcement mentioned that it's less cost competitive per task than Opus at higher thinking levels. It's significantly cheaper at lower levels though. I'm wondering if this is going to be a universal pattern of smaller models: they're less smart, so to achieve the same benchmark results they have to think a lot more and hence become expensive. Benchmarks force models to solve the problem entirely by themselves, requiring thinking. But if you pair them with a smart model (who thinks and solves beforehand) they won't need to solve the hard parts and can run on low/med. I suspect that was Anthropic's intention.
- pier25 3mo agoOn top of that isn't it strange that if the LLM makes a mistake you're still charged for those tokens? They're selling "intelligence", automation, etc but if the service doesn't work as expected the user has to pay for that.
- wccrawford 3mo agoYou're paying for a service with known flaws. They do not guarantee correct answers. Also, LLMs don't "make mistakes". They don't think or act. Every single thing they output is a hallucination. It just happens that the vast majority of things align with reality.
- Aeolun 3mo agoIf I use electricity to do something stupid, I still have to pay for the electricity. Intelligence is just another utility.
- pier25 3mo agoSpend 5 minutes on the marketing pages of any AI company and it's obvious they're not selling electricity, fuel, or even raw compute.
- SwellJoe 3mo agoEfficiency is the next frontier in LLMs, and I'm not confident the American companies are taking it seriously enough. DeepSeek, even in a naive API-calling loop, serves something like 80-90% cached tokens at an absurdly low price per token. Using an agent harness tuned specifically for their caching (Reasonix) pushes the cached tokens to 97-99%. DeepSeek is consistently among the cheapest models per task in my benchmarks, while also performing quite well. I'm still almost always using Claude for work, but for side projects, small stuff, etc. and anything better served by an API rather than starting up Claude Code (or `claude -p`) I'm using DeepSeek pretty often. Anthropic models also shut down on a lot of security-related work, which is what I've been spending a lot of time on lately. I expected Fable to refuse this kind of task, but even Opus 4.8 refuses to build a verification harness for security bugs, as that involves exercising a discovered bug to prove it's been fixed in an automated red/green way, which looks like exploit creation to Opus' guardrails. So, I have to use other models for that work, now, though most of the original benchmarks I built were built with Claude.
- cyanydeez 3mo agolocally, im about 80k every 30 minutes for a project. can run deer flow to pump them tokens. but yeah, its the 80s LOC metric since quality isnt captured
- yalogin 3mo agoAnother thing I noticed is the llms perform efficiently/effectively only under the optimum circumstances. May be this just a Claude issue, but when the session goes on for very long the effectiveness drops drastically and I start getting bad answers. This is especially true for design and debugging. Wonder how that ladders up to the token usage
- gkrishna 3mo agoAgreed with the problem, but what's the solution?
- dualdust 3mo ago[flagged]
- thraxil 3mo agoNot every application uses LLMs the same way. For some use-cases, price per 1M tokens is absolutely meaningful. Eg, we do a lot of pretty basic classification/entity-extraction/summarization type work on large inputs (100k tokens per request being very common). It's pretty easy stuff; Gemini 2.0 Flash was perfectly adequate, quite fast, and cost $0.10 per 1M tokens (and even less when we could make use of the batch API). Every newer more powerful model obviously can handle the same work but costs significantly more. When we're deciding what model to use, price per 1M tokens is definitely a meaningful metric.
- chongwei89 3mo agoDo we have any good benchmarking for "cost per tasks done" vs "cost per tokens" comparison across models?
- gawkdev 3mo agothe meaningless part is right but doesnt go far enough.. price per token also assumes every token does the same amount of work.. a model that reasons for 10x longer to get the same answer is not actually cheaper just because its rate is lower. the number that actually tells you something is cost to get a task done end to end.. not the sticker price on a token. anyone tracking that instead of the headline rate ??