9 ms·
The unbearable cheapness of open weight models
- linzhangrun 4mo agoIt would not be surprising if GPT and Claude get cheaper too as inference gets cheaper. Two years ago, o1 was the strongest model and cost much more than Fable, while being nowhere near as smart as a Qwen 3.6 35B that you can now run on a DGX Spark without much trouble.
- ddxv 4mo agoTrue, outside of the dark tactics I imagined in the article, they will have to compete at lower costs. It's just that the current iteration does not feel cost competitive yet.
- tsss 4mo agoProbably they will, unless Claude and GPT become luxury brands like Gucci. Currently it makes no sense for them to invest into efficiency. They need to put everything into competing for the top spot as long as they still have a shot.
- an0malous 4mo ago> It would not be surprising if GPT and Claude get cheaper too as inference gets cheaper No because the biggest factor in their current price is VC subsidization which has likely peaked if OpenAI is now serving ads and Anthropic has increased their API pricing
- deleted 4mo ago[deleted]
- odie5533 4mo agoThis is what concerns me about how AI giants are planning to make money. Their product has already been commoditized at prices which for them are still subsidized to grab market share. Unless the giants invent a technological leap, their prices are going to be dragged down by open weight models and I don't see how they'll turn a profit.
- Jimega36 4mo agoReach AGI to leapfrog whoever is behind. Burn everything to get there faster.
- jorisw 4mo ago'Reach AGI', the same way SpaceX will put data centers in orbit. A pipe dream.
- ben_w 4mo agoI'm currently writing a blog post about data centres in orbit, and my current conclusion is that even though they can build one, they definitely can't put 1 million up there and would have better things to do if they could. AGI? Too loosely defined. They lack a lot of competences which humans recognise when we see them but find it hard to put into words; on the other hand what they can do they already do faster than any human (and have greater breadth than any single human, but this usually doesn't matter because "coder" and "economist" and "translator" gets solved in human teams by hiring three people). I do not think current ML has the tools to solve for quality. But we know it's possible for a really mediocre intelligence to make human level intelligence, because evolution made us, so for me the question of AGI is more a practical one: is it affordable? (I also think not at the present time, but that's an "I think" not "I am analyzing it carefully").
- trick-or-treat 4mo agoMaybe you missed the part where starlink / orbiting datacenters don't really have to even make money as long as they partially fund rocket launch tests. Or maybe you don't take Elon seriously when he talks about Mars.
- ben_w 4mo ago> Maybe you missed the part where starlink / orbiting datacenters don't really have to even make money as long as they partially fund rocket launch tests. I am only dismissing the orbital data centres, I do see a future for Starlink. One with competition, but a future nonetheless. I'm old enough to remember the dot.com bubble and "we lose money on each unit and make up for it in scale": If they don't make sense, they don't help. Putting a single one in space, or even a handful, is physically possible! But even optimistic Alphabet researchers (and Alphabet owns more of SpaceX than the entire IPO) say this only makes sense at $200/kg, while early Starship launch costs while they sort out reusability be at best $400/kg and the researchers don't expect $200/kg until the mid-2030s even with a high launch rate: If the learning rate is sustained—which would require∼180 Starship launches/year—launch prices could fall to <$200/kg by∼2035 - section 2.4, https://arxiv.org/abs/2511.19468 https://arxiv.org/abs/2511.19468 At $200/kg, and using the payload estimates elsewhere in the paper (the learning rate is based on mass rather than launch count), they'd need to launch 370,000 tons (4.4 ibid); even at the "good enough" cost, $200/kg, they'd need to spend $200/kg * 3.7e8 kg = $7.4e10. That's a hell of an R&D spend for the next 10 years of a company whose lifetime revenue (not profit) is reportedly $4.6e10. My current draft has a few thousand words of additional problems, plus a bunch of things which I mention only to say why they are not, and some more where I say the research has yet to be done. > Or maybe you don't take Elon seriously when he talks about Mars. Used to, not any more. Has been too slow with Starship even before the fact that iteration with hardware is necessarily slowed down by a 2-year gap between launch windows. There's not even been any news about demonstration models of either Mars-rated or Starship-rated Sabatier processors, which would be an easy win and also win points for both environmentalism and energy independence viz. Iran/Hormuz.
- arikrahman 4mo agoWith cache hit rates being effectively free, harnesses like Reasonix have let me do a month of work for less than 2 dollars. It's not even the subsidies making it cheap, American providers like Digital Ocean or Cloudflare host the same model with similar pricing.
- ForHackernews 4mo agoI think this is very likely and something that everyone seems to be missing when valuing these AI firms. AI is not the new industrial revolution, it's the new cloud VM: a very useful commodity software offering.
- antonvs 4mo agoThe parallels to the Industrial Revolution are so close that we even have a new generation of Luddites. (Not saying they don’t have some valid points; so did the original group.) The reason it’s like the Industrial Revolution is simply that there’s no question it’s going to completely transform jobs. It can make a very similar difference to the difference between a craftsman and a factory worker. The latter is massively more productive.
- Scaevolus 4mo agoCloudflare's Deepseek V4 Pro prices are 4x more than Deepseek's for input and output tokens, and 100x more for cached input tokens, which is crucial for the tool uses of agents which cause multi-turn conversations.
- arikrahman 4mo agoCache hit is less than a cent with Deepseek Flash and 3 cents with Cloudflare, it's free vs almost free. Where are you finding the statistics on Deepseek Pro? I don't see Cloudflare as a provider on openrouter for Pro, only flash.
- pjc50 4mo agoHow does caching help here? How much repetition is there in queries?
- Jackobrien 4mo agoThe giants knew this was coming, and soon 95% of AI tasks will be able to be done by open models (coding, research, cowork style work). So why pay a premium? Why use them at all? This leaves the labs with two options: 1) push the frontier in a way only massive scale can, and cash in on it (mythos level cyber security, recursive training, frontier science work). There’s big money for never before possible capabilities. 2) own the app layer with their edge in reputation and powered by their infrastructure. Be apple where everyone else is Linux. Do design, coding, research, SMBs, legal, finance, healthcare and more (they are doing all of this). Will it be enough to justify a Google level valuation? We’ll see how fast they can push it.
- ed_elliott_asc 4mo agoWon’t all they need to do is say “best in class, latest models, fastest” and wine and dine a few execs and those enterprise deals will be signed? In this case the people tasked with using the product won’t actually mind.
- actionfromafar 4mo agoYes, exactly that. Be Azure and Office 365 and Sharepoint and AWS where everyone else is Debian Stable on a USB thumbdrive.
- fragmede 4mo agoOffice 365? Ew, Google docs, please.
- NitpickLawyer 4mo agoNo one is getting fired for using SotA.
- spwa4 4mo agoIf the price difference is 2x? Sure. If the price difference is 50x? No way.
- arthurofbabylon 4mo agoLet's imagine that Anthropic/OpenAI fail to manufacture scarcity by villainizing Open Weight models (a sincere probability). What is left for these corporations to prop up their prices, or any margin at all? I expect scaffolding around tool use, supporting bespoke implementation and driving risk down for institutional adoption. (They might even build an insurance tool to protect accountants/lawyers from errors in compounded probabilism!) A question for economists... It seems plainly clear to me that information and information processing is commodifying (for the first time in human history?). Without the age-old bottlenecks at the top of the value chain, capital will surely flow downwards, right?
- ddxv 4mo agoOpenAI, though they seem to backtrack it lately, have been slowly pushing forward of their launch of ads which would be a supplemental way to support cheaper use of their models. This is currently not as great a fit as the modern day banner ads, but it will be interesting to see where they go with that.
- AnthonyMouse 4mo ago> It seems plainly clear to me that information and information processing is commodifying (for the first time in human history?). Without the age-old bottlenecks at the top of the value chain, capital will surely flow downwards, right? Isn't this the thing people have said about every new technology since the printing press? And it has been mostly true, but it has also been the case that the incumbents have fought hard to lock things back up again. Newspapers and radio stations buy each other up, the open web gets locked inside Facebook (which, 30 years ago, people were already worried about with AOL), people have computers in their pockets they can't run their own programs on anymore. Interests are going to want to lock the new information thing behind a gate so they can charge a toll and censor what they don't like, same as it ever was. You don't win by default, you have to fight to stop them.
- arthurofbabylon 4mo agoI don’t think that comparing LLM’s to the printing press (and radio, film, TV, etc) is an apt analogy, and I don’t think that people have said the same things about the two technologies; the prior technological changes in information dealt with distribution, while this one deals with processing and production. Recall the notion of a bottleneck, and this distinction will become clear. Those prior technological changes never inverted a bottleneck, and this one does.
- surgical_fire 4mo agoOne thing it doesn't even mention is how good those models are. Evet since I moved to DeepSeek I had zero regrets. It performs exceptionally well. I honestly prefer it to ChatGPT (or Claude that I use at work). I never used Fable, maybe it is that much better. DeepSeek has no problems with the workloads I give it though - if it only keeps marginally improving with each interaction I don't see myself needing to come back.
- my-next-account 4mo agoI wonder whether Oracle is going to go bankrupt because of this
- worldsayshi 4mo agoWhy Oracle?
- InsideOutSanta 4mo agoThey're extremely exposed to a market crash due to their huge debt-funded compute contracts. Having said that, while one can always hope, I would assume that Oracle is one of these companies that will be bailed out or find a way to survive.
- cyanydeez 4mo agooracle is licking so much boot, you'd need to also have the republican fascist party completely faall apparent.
- anax32 4mo agoOpen weight and local hosting is far, far cheaper. In every respect. Even support is cheaper, over time. However, it's difficult to sell this to businesses who want contracts and KPIs, not staff and commitments. Regulated industries will favour the closed sources, either by choice or mandate. The interesting question is whether they will have better models, or worse models. History says they will receive a worse service, but continue anyway.
- general1465 4mo ago> Regulated industries will favour the closed sources, either by choice or mandate Until your country will appear on naughty list of US administration because your local politician did something what mildly inconvenienced US oligarch
- tuatoru 4mo agoCheaper until you factor in security and liability, which are going to get increasingly salient over time.
- dist-epoch 4mo agoIt's so refreshing to read a short to the point article, which is not extruded into 10 pages with LLMs.
- leroman 4mo agoThe token-economics for closed source models are different, they are optimizing for 200 USD tokens worth of software engineer monthly usage, they will increase per token price as models or harnesses are more optimized.
- isoprophlex 4mo agoAren't these open models so cheap because they're (partially) chinese gov. sponsored, and because they're stealing and redistributing the IP that comes in?
- grebc 4mo agoAnd the American ones are stealing and redistributing the IP of every single person who authored anything on the internet at some point.
- blamestross 4mo agoWell I can't speak to the chinese gov part, but ALL the models are IP laundering systems. I'd rather IP get laundered into open source.
- jrm4 4mo agoTechnically correct, the worst kind of correct :)
- titanomachy 4mo agoMaybe, but there's tons of providers available, so you can pick one that you trust not to steal your IP (or run it yourself, if you're rich and paranoid enough).
- amanaplanacanal 4mo agoWhose IP do you think they are stealing? According to US courts, training is fair use. And even if it wasn't, they are distilling output from other models, which isn't copyrightable, again according to US courts.
- danny_codes 3mo agoHow dare they steal what we already stole!
- nsoonhui 4mo ago[flagged]
- bmnbmnbmn 4mo agoOne of the purposes of open weight models is to create a moat. If there were no open models available, I think we'd see much more and better models coming from Europe by now. Right now, any startup wanting to build and sell a model needs to be substantially better than the open models, which has become increasingly difficult and expensive.
- tuatoru 4mo agoEurope has Mistral. You and readers may be interested in Europe 2031 1. https://europe2031.ai/ https://europe2031.ai/
- snootypoot 4mo agoi agree with his statement that the big companies and the string pullers in government are inching toward banning open models.allowing the plebs unrestricted access to things seems against the wishes of the "you will own nothing and be happy" / "you will rent everything on the cloud and subscribe to your appliances" crowd such as blackrock and so on. anyone who disagrees is not seeing the forest, only the trees.
- amanaplanacanal 4mo agoI don't see how they could ban them in the US. Code is speech, and the first amendment still mostly holds. They might try, but I don't see the courts upholding it.
- juancn 4mo ago90% of my model use is on local open-weights models. The things that I need to automate do not need frontier models. Heck, even a gemma-4-12B-it-qat-UD-Q4_K_XL can deal with a lot of complexity if properly guided (it can run on 16GB of unified memory, for example on a base model Macbook Air). I've been using it to translate Javascript to a custom scripting language in a product I work for, just by providing a system prompt and an MCP tool to call the target compiler to check for errors. Sometimes it converges faster than Opus 4.6 (I've tried) because it doesn't over-think stuff. If it were a person I would say it knows less, but it's still smart. I mean, you don't need the most powerful tool at all times. We treat AI as one-size-fits-all, and once cost gets in the way, it will matter.
- beepdyboop 4mo agoI don’t get it. So many here are saying open weight models will kill the frontier labs. But open source and similar have tried to beat private companies everywhere all the time, and people still buy the best products even if great open source alternatives are available. Why wouldn’t this be the case for AI too?
- drudolph914 4mo agoI feel like this comment is just engagement farming, but I'll bite anyways there is a larger appetite for something like open source AI mostly b/c of price. we all know these labs have not figured out their pricing model, and we're all holding our breath out of fear of what the prices could be. also, if you consider that the only toll to knowledge work before was personal time, and now you need to pay $100s month just to keep up with the baseline speed. it makes sense people are looking for something that gets them back to a workflow where the price to do work is near $0.00. I think for a smaller group though, it's more to do with a certain combination of principles. Some people don't want censorship, other's want ownership, some want the knowledge of working on LLMs to not be gate kept.
- dominotw 4mo agoppl buy iphones over cheap android phones . android phones can do everything that an iphone does ( and better)
- 05 4mo agoWhich cheap (or expensive) phone does better log/raw video than iPhone 17 pro max?
- mxschumacher 3mo agothere's still some lock-in here: data storage (iCloud), the broader Apple ecosystem, certain apps, user habits (I'm a lifelong Android user and am having trouble using other people's iPhones)
- an0malous 4mo ago
- kittikitti 4mo agoEven if open weight models were vastly more expensive, I would still prefer them. I don't know where my data is going and whether they're lying about the model when I make an API call. They can ban you from their API for any reason. Anthropic recently pulled their frontier models. There are numerous compliance concerns. The list goes on and on.
- dwaite 4mo agoWhere's a few good places to go to learn more about open weight models, both running hosted and running locally?
- drillsteps5 4mo agoAside from googling "how to download and run open weights model" check out localllama (yes 3Ls) subreddit. Huggingface.co is where many of them are published. There's many providers that run open weights models and give you access. Many decent open weights models cannot be run on consumer-grade hardware (DeepSeek, GLM, many others).
- macwhisperer 3mo agoliterally ask cloud ai like the free gemini or chatgpt.. they could make you an expert on the subject overnight..
- CuriouslyC 4mo agoThe government is going to ban foreign models and foreign inference providers, without question. The US govt is going to dig its dirty little fingers into OAI/Anthropic/Oracle/(probably)SpaceX and end up taking some stock for a sovereign wealth fund (probably timed to prop up flagging share prices, and with the promise of sweet government grift down the line), and at that point the bans will be framed as protecting that investment.
- Zak 4mo agoOne issue I keep seeing with cost comparisons is that they compare API rates while a substantial fraction of users are on subscription plans. It's more expensive to use GLM 5.2 paying z.ai or Opencode Zen API rates than it is to use Opus on a subscription plan. Both of those providers offer subscriptions priced favorably relative to their API rates, but only in what are effectively trial sizes.
- 1matin 4mo agoAnd that means either: 1. They overprice their APIs to make their subscriptions look reasonable 2. They burn money with their subscriptions
- Zak 4mo agoCould be a little of each, plus a third option: subscription users don't always consume their entire quota.
- cherryteastain 4mo agoEnterprise plans don't have the equivalent of the subsidized-usage-included Claude Max/ChatGPT Pro plans anymore. The revenue generated and total amount of tokens used by individuals is probably a tiny fraction of tokens billed at API pricing.
- tuatoru 4mo agoDeepseek's price looks unsustainable. Ant have said their operating margin is 70%. A leaner company could maybe raise that to 90%. Most of the cost of supplying inference compute is depreciation of the GPUs. Maybe Deepseek is anticipating a 50 year life for theirs.
- Tuna-Fish 4mo ago> What worries me about this is that Anthropic and OpenAI seem to have backed themselves into a corner of high costs. Can they reasonably decrease their prices by 20-50x to compete with DeepSeek or Xiaomi’s Mimo? They have high prices, not high costs. They will obviously keep prices as high as they can for as long as they can, while keeping demand up. Once demand starts to fall, so will the prices. > Are these models cheap because they are open weight and having hundreds or people stress test running them on different hardware helped to lower the cost? Or is it that they are being provided as loss leaders to drive the prices down? Neither. They are cheap because they have neither technical edge nor brand power to keep the prices high, and so have to ask commodity prices for them. People somehow still don't get it, despite everyone who studies the economics of it telling them: Inference is dirt cheap. Training is expensive, inference is cheap, and getting cheaper.
- Schiendelman 4mo agoI get it! And I appreciate people like you pointing out the business side of LLMs. Also, these open weight models are significantly lower quality than the high end coding models, and for some reason a lot of people think they're exactly the same. Maybe engineers who only dabble in LLM usage aren't doing enough complex work to notice...?
- manwithopinions 4mo agoSo why are they losing so much money? Money is made on the subset of inference that is charged at cost + margin via their APIs. API usage is so high because customers are still finding their feet, trying to understand how to measure the value they get from their spend, erring on the side of spend. Yes, in a world of unmeasured value and tokenmaxxing, inference is profitable on SOTA models because all capacity is being consumed at all times, driving down marginal costs, but what about a world in which capacity isn’t constrained? There are still huge fixed costs. Even the most optimistic leaks with the current high prices put the margin on API token inference at around 50%. How can SOTA models ever come close to competing on price? Price always matters. Offering the best model with the most brand recognition does not exempt OpenAI from the basic rules of business. Historically, software has been such a successful business because the margins are incredible, 95%+ in many cases, driven by direct measurable value to customers that dwarfs the cost. A 50% margin at a time when your customers are falling over themselves to spend as much money as they can is not a good sign, it is a very bad sign, it leaves no room to ever achieve traditional technology margins, and inevitably leads to very weak margins. Inference needs to become an order of magnitude cheaper than the value it delivers to ever have a chance of delivering on this wildly profitable vision. The cheap model providers have a much better chance of achieving that. Outside of coding, almost every business case for AI doesn’t need above human intelligence, it doesn’t even need human intelligence, or half a human intelligence, a business can extract a lot of value from a machine that has a fraction of a human’s intelligence. Most human work does not use our intelligence, it is rote, a monkey could do it, and that’s where AI will be used most. Who is going to pay $10 per million tokens when they could pay $0.10 to get the same outcomes?
- MoonWalk 4mo agoI'd appreciate an explanation of what "open weight model" means. Is it a "weight model" that is open, or a model with open weights (so should be "open-weight model"), or is it weights that can be applied to a model? Are weights separable from a model? And if not, what is the point of saying "open-weight model" instead of just "open model?" To the newcomer, it's hard to determine what the components of an AI system are from the throwing-around of these terms.
- philipkglass 4mo agoA completely open model is one like the Allen Institute's Olmo model series: https://allenai.org/olmo https://allenai.org/olmo The trained weights are open, the training software is open, and the data that goes into training the model is open. Not many models are fully open. An open weights model is one that has freely available trained weights, and maybe fine-tuning tools, but it lacks the original training data (and usually lacks the training software). These are the most commonly used local models, like Google's Gemma series, Meta's Llama, or Alibaba's Qwen.
- MoonWalk 4mo agoSo you can apply different weights to those "non-open" models? Also, I've read a bunch of descriptions of AI components, but none of them has said what the weights are applied to in the model. I guess that every model contains a dictionary of words and phrases, and the weights map relationships between them? All the descriptions simply talk about weights being applied to "input," but neglect to say what that input is compared to. If a user submits a query, are the words in the query weighed against the words in the model? Can you recommend a primer on this whole process?
- philipkglass 4mo agoThe weights are just numbers. I don't what technical background you have in other areas of computing, but I think that this is a good, short introduction that doesn't assume too much: https://www.3blue1brown.com/lessons/mini-llm/ https://www.3blue1brown.com/lessons/mini-llm/ To quote part of it, Training a model can be thought of as tuning the dials on a really big machine. The way that a language model behaves is entirely determined by these many different continuous values, usually called parameters or weights. Longer and slightly more technical, "Intro to Large Language Models" by Andrej Karpathy: https://www.youtube.com/watch?v=zjkBMFhNj_g https://www.youtube.com/watch?v=zjkBMFhNj_g
- cws_ai_buddy 4mo ago[flagged]
- drillsteps5 4mo agoOpen weights models are cheap in the context of the article (when you run inference in the cloud) because they are free. When I pay for inference for running DeepSeek open weights model I only pay the inference service provider for compute/memory/storage/network throughput. The model itself is free, the developer isn't getting a dime. Developing these things is NOT free, there's a lot of labor, hardware, compute/memory/storage/network that goes into that. Who's paying for all this? Chinese govt? Developers themselves? What's the revenue model here? I absolutely LOVE ability to either run them locally or access inference providers on the cheap, but having a hard time understanding the financial side of this.
- linzhangrun 3mo agoAnother thing these AI giants should worry about is the local deployment A $4,699 DGX Spark can easily run Qwen 3.6 35B, and its performance crushes GPT-o1, the SOTA from two years ago
- cws_ai_buddy 3mo ago[flagged]