13 ms·
Open-weight AI is having its Kubernetes moment
- firasd 2mo agoOne of the strangest things in the AI industry is 'tokenomics'. It's not very clear why using GPT-4 in early 2023 was so expensive and then six months later 20 bucks could get you a fair amount of GPT-4 inference. This pattern has continued across various labs/providers for years--there is a continuous see-saw of pricing that doesn't seem related to anything. So what open weight models do is at least provide a baseline of inference cost to add some sanity to the price markers. And of course predictability too--if you really want Kimi K2 instead of K3 you can still use it. So the competitive pressure and predictability offered by open models is helpful for users
- minimaxir 2mo ago> It's not very clear why using GPT-4 in early 2023 was so expensive and then six months later 20 bucks could get you a fair amount of GPT-4 inference. FlashAttention was a hell of a drug.
- mawadev 2mo agoIts very clear: nobody wanted to pay for usage at that price point
- serial_dev 2mo agoIt's basically supply and demand?
- julianlam 2mo agoWhy do prescription medications cost so much, and generics so little (comparatively)? Artificial inflation to recoup R&D.
- gwbrooks 2mo agoNot sure pricing to recoup costs is artificial.
- ncallaway 2mo agoThe government backed monopoly to ensure that supply remains artificially restricted to ensure that the market will support the higher prices is
- andsoitis 2mo agoDrug companies have a portfolio of compounds they research. Most don’t pay off, so R&D costs make their way into the pricing of those superstar and other drugs that do work. Also, timelines are pretty long.
- mjhay 2mo agoDrug companies spend more on marketing than R&D
- andsoitis 2mo agoSo?
- DennisP 2mo agoSo pharmaceutical companies spend far less on marketing outside the US, partly because every other country besides New Zealand makes those incessant drug ads illegal, and partly because governments negotiate prices and keep profit margins down. If the argument is that R&D costs are what make drugs expensive, then we could easily eliminate an even greater expense by just copying what other developed nations do.
- spwa4 2mo agoBecause as per usual it's silicon valley misunderstanding economics. AI is HPC. And how the HPC market worked before: If you're the best performing "computing cluster" (ie. whatever you call the entity that can complete a massive calculation), you get a blank check from Congress. Why? Because you need those calculations to "pump" nuclear weapons. They are needed to calculate both the geometry to make fusion bombs possible at all and to calculate the effect of a given geometry. They are the reason US/Russia/China have the biggest and strongest weapons known to humanity. And of course, they were replicated worldwide for this reason. I mean not that anyone will admit this but we don't have the best possible solution, and we don't know either the upper or lower limits for fusion devices (plus the lower limit would be very useful for energy generation, which for the US would effectively mean almost literally unlimited large marine ships that never need refueling. And yes, the solution to that problem is almost literally a 3d shape. Not just that, but mostly) Now it appears it does not work the same when you democratize computation. Humans want a particular amount of computation and are willing to pay a given price for that. But the accountants still saw the blank check from before and ... do what accountants do. Economics don't change because you make things bigger and accessible, do they? Oh ... wait a second ... As someone put it recently though, we now have data. 2.3% of humans in the US are willing to pay $20 per month for the support of a model like GPT-5.5/Claude code. If that's true (and after years of having this model, why wouldn't it be?) ... it means AI startups are doomed (because it's not even 10% of what they need it to be to make economic sense).
- gwbrooks 2mo agoI disagree with some of the framing, but that's what good discussion is about -- figuring out where we agree and disagree. But consumer uptake strikes me as the worst way to judge whether the big AI shops will make it. That's not where most of the leveraged user return or deployable capital is.
- m4rtink 2mo agoI'm sure Teller could make you a 10 gigaton nuke with a slide rule if you did not mind some sub scale tests.
- mlyle 2mo ago
- thewebguyd 2mo ago> if you really want Kimi K2 instead of K3 you can still use it. I think this is a very important aspect, especially after the huge GPT-4o backlash when GTP-5 came out. Each model has certain quirks, and areas where the previous model might be better for some use cases than the latest and greatest, and the labs so far seem to have no desire to offer some kind of "LTS" release.
- jmyeet 2mo agoYeah I've been thinking about this and the analogy I came up with is that tokens are basically equivalent to an in-game currency in -free to play" mobile games. You're trading actual money for some notional "curency" or coins that can only be used for one thing but, unlike mobile game coins, you don't know how many coins something costs before you use them. It's kinda weird.
- bee_rider 2mo agoI think it is worse actually. Tokens in a game are usually just purchased for enjoyment in the game. They are purchased as a part of your entertainment budget, not expected to be useful in any way. LLMs are fundamentally tools intended to be useful. But LLM vendors don’t understand their systems well enough to actually price the product people are trying to buy (for example, the actual product of a coding model is the code that it produces, not the tokens, which are just an internal mechanical process involved in the creation of the code). Token based pricing is that lack of understanding leaking out of the organization that ought to be responsible for it, and being dropped on the user. Imagine if we made cars like this! You’d go to the car dealer and ask for a car. They’d bring you a pile of parts, charge you for them, and try to put them together in front of you. You’d go back and forth for a bit, rephrase where you want the steering wheel, etc. Some of the parts wouldn’t fit but you’d be invited to pay for replacements as well. In the end you’d either have a car or not, that’s your problem.
- vikramkr 2mo agoDid they actually ever cut the price on gpt 4? The oldest versions of it in the api still seem stupidly expensive? There were definitely price cuts as they introduced the turbo models and stuff, and new versions of each model might have gotten pricey cuts, but just because they're both called "gpt-4 something" doesn't mean they're the same under the hood or that they didn't change a bunch of stuff under the hood to make it cheaper to serve
- esseph 2mo ago> because they're both called "gpt-4 something More like a generation of models with different specific use cases
- simianwords 2mo agoWhy is this so difficult to understand? 1. the field was nascent and new efficiencies were discovered 2. supply and demand 3. its in the company's incentive to make their models more efficient to increase overall usage so that while the margin remains the same, the total revenue + profit increases I genuinely don't know what puzzles everyone?
- Aurornis 2mo ago> It's not very clear why using GPT-4 in early 2023 was so expensive and then six months later 20 bucks could get you a fair amount of GPT-4 inference. The price is what the market is willing to bear for the available compute capacity and competitive landscape. You can only discover that price after trying different price points and seeing what happens. Everyone is trying different pricing schemes and discounts as they test the market. The demand is fluctuating at the same time. It’s probably very confusing if you’re primarily familiar with stable and mature markets. Price fluctuations are a common feature of new and evolving markets.
- imachine1980_ 2mo agoMost unmature markets aren't subsidized to the point that LLM market is, most market have some level of baseline profitablity, this market doesn't, that's because most market subsidized the marketing or the capex but this market doesn't hold the opex, the capex not the amount of marketing let alone all of this together
- Aurornis 2mo agoMost new markets are funded by initial investment capital. Early entrants operate at a loss as they grow. This isn’t as unusual as some people are trying to make it sound. This has been happening since the dawn of finance. I thought this would be less foreign to everyone since we just went through this whole conversation for a decade with Uber and Lyft. Their demise was predicted from the start from everyone who thought that it was going to collapse as soon as they couldn’t subsidize your rides with promos. There was much wailing and gnashing of teeth as their prices changed to feel out the market. Then they found profitability and the critics went silent.
- PunchyHamster 2mo agoArguably we'd be much better off if none of those would be subsidized by investments, at least not to the "run unprofitable for decade+" level. Because that just absolutely murders any competition that manages to not get that level of free money. You're not pouring money in to make it happen at all at that point, you are pouring money in so nobody else can get the part of the pie. Which is great for investors, bad for everyone else
- mountainriver 2mo agoQuantization also started picking up around then, as well as distillation into smaller models
- RobRivera 2mo agoMarket discovery
- jrm4 2mo agoIt's only strange if you do the silly thing of presuming a "fair market" in which e.g. it's generally easy to get reliable information about how all of the things work. There's just obvious and enormous incentive for the OpenAI's of the world, along with all of the other players, to confuse, misrepresent or just straight up lie a whole bunch about everything given how new and unknown the tech is.
- segmondy 2mo agoWhat is strange about GPT-4 being expensive in 2023? Supply and demand. Which other model choices did we have? Prices are related to supply and demand. We see it play out with the introduction of capable open weight models or even other closed cloud models.
- firasd 2mo agoNot really though right? As of early June 2026, Opus 4.8 in fast mode cost $50/M output tokens and Opus 4.6 & 4.7 cost $150/M output tokens in fast mode How can supply and demand explain the price drop? Was it cheaper to serve Opus 4.8? Is the demand for the newer Opus lower than for the older Opus? These are just fixed prices that seem picked out of thin air
- hluska 2mo agoThey would have been picked out of thin air. That’s the joy of innovation - you have to randomly throw prices against the wall and see what sticks. The point where it sticks might be equilibrium or it may be an inefficient market… and nobody will know which one until it’s too late.
- hluska 2mo agoThat’s a slightly naive view on pricing. That equilibrium point doesn’t just magically appear - it’s found through price testing.
- cr125rider 2mo agoWay over complicated for what most users need? Huh?
- deleted 2mo ago[deleted]
- netdur 2mo agowhy would any software want to have Kubernetes moment? can't count how devop I know that is confused by it
- xyzsparetimexyz 2mo agoI still don't know what it is tbh. Something for docker?
- spicyusername 2mo agoYou... don't know what Kubernetes is... pretty impressive, honestly. Its 2026 and its the de facto method of deploying software basically everywhere. You gotta really work for it to not know what its for by now.
- chrisandchris 2mo ago> Its 2026 and its the de facto method of deploying software basically everywhere. That is some really impressive bubble you are living within. Basically everywhere - nowhere near that, no. [edit]: Maybe containers, but software in general is so much more broad than containers.
- recursive 2mo agoI don't do much deployment but I'm in the same boat. Something something docker automation?
- YetAnotherNick 2mo ago2026 is the year for vercel and Render.
- RussianCow 2mo agoI haven't used them but aren't they basically modernized Heroku? What's different?
- curious_cat_163 2mo ago> The government should use procurement to create demand for portable, interoperable systems rather than permanent dependence on one API vendor. Now, here is an idea that I have not heard before... and I think there is some merit to this. This is also the sort of thing that a state (looking at you CA, CO, IL, NY) could do, instead of just the federal government.
- tangotaylor 2mo agoI dunno about the other states I feel like we need a better coalition lobbying California lawmakers about AI innovation because I feel like they're doing a lackluster job. For example, Buffy Wicks introduced AB 2023 (passed the Assembly) which will effectively cause chatbot operators to ban minors because of the huge liability risk that the law introduces. This is a great way to kneecap innovation by excluding curious children and teens who are often the most innovative. California already has laws on the books to address chatbots encouraging self harm (BPC §§ 22601-22606) so I don't know why on earth they're doing this. Then there's AB 2169 introduced by Lowenthal, which would have mandated interoperability between chatbot platforms to help people migrate to competing ones easier. I thought this was awesome but it didn't even get a vote in the Assembly. Maybe the newly-created Little Tech Association can help here.
- nttylock 2mo ago[flagged]
- thih9 2mo agoIs anyone using open weight models for agentic coding? What is your stack (harness, model) and how much do you pay per month? How would you compare your experience to a typical subsidized plan like Claude Code + Pro plan? I’m asking because i keep hearing that open weight models are cheap and efficient - is that really the case in practice?
- rglullis 2mo agoI am using GLM-5.2 via Ollama Cloud in the $20/month plan. With the same plan I can get many different API keys that I use to run my OpenWebUI server, my opencode and pi dev sessions. I am usually running 2 to 4 sessions concurrently, and I never hit quota limits. At work I get Claude, and I was getting reports that I was spending $75 per hour of work on Opus.
- revolvingthrow 2mo agoWhile I am grateful for open weights models I never found much use of them in the past, barring those I could run myself. This changed with deepseek 4 - it is staggeringly cheap, even if the performance definitely isn't near sota and it's not particularly fast either. When I expect to need a lot of tokens and the task isn't too difficult I use sota to plan and create a thorough set of instructions and let deepseek chip away at it. With thorough instructions the quality tends to be satisfactory, and you pay something silly like $15 for 600m tokens. GLM 5.2 seems like a decent price/perf and Kimi 3 has some real nice performance for an open weights model, but gpt 5.6 is unexpectedly affordable (especially if you don't automatically use Sol at max) so I don't think either is worth it atm. The exception is when you're working on something that US models get cold feet about, which seems like a constantly growing list. For me Fable is already too much of a headache in this regard, but chatgpt is still okay-ish. Hopefully it'll last. If not, there's Kimi. tldr SOTA for most things because gpt 5.6 is token efficient. If I expect to burn a lot of tokens I use deepseek 4.
- Scene_Cast2 2mo agoI'm using Kimi K3 + OpenCode. I pay their API pricing, costs about $5 / hour (and chews through ~10 million tokens / hour) during continuous use when I have one or two sessions running and doing their thing. Can't comment on how it compares to plans (I really don't like the limitations and general shenanigans I see around plans, so I've never tried them). It is notably slower than Fable / Opus / Gemini, but also vastly cheaper than their API pricing.
- deleted 2mo ago[deleted]
- chrisjbg 2mo agoDario is a FUD-spreading douche
- claw-el 2mo agoEssentially doing this, https://health.clevelandclinic.org/catastrophizing https://health.clevelandclinic.org/catastrophizing while making other’s mental health worse..
- YetAnotherNick 2mo agoIn fact more countries should have government funded models. There are some obvious issues in China completely dominating open weights space. Kimi had funding of just $2B and could literally create national security threat. A lot of countries could fund something in the range of few billion for something so important. At the very least US and EU could fund few companies.
- debarshri 2mo agoShameless plugin. Funny enough we just made agents kubernetes native at adaptive [1] [1] https://adaptive.live https://adaptive.live
- petilon 2mo agoEnormous amounts of money is being invested in the development of AI models. Investors expect returns on their investment or they will not continue investing. Open weights make it harder for investors to get their money back, so it harms the industry. Once the weights are out, it makes no sense to ban them in the US while the rest of the world takes advantage of it. But that doesn't mean developers of frontier models shouldn't take steps to prevent their weights from being stolen.
- jumpkick 2mo agoHow are weights stolen from the frontier model developers? What is that actual mechanism?
- eigenspace 2mo agoThe weights themselves aren't stolen. The claim is that Chinese companies are using VPNs and proxies to buy massive amounts of Claude Pro and Codex subscription accounts, and then selling usage on those subscriptions as cheap white-label LLM API usage. While selling that LLM API usage, they then capture all the prompts, outputs, and intermediate thinking the LLM does, and then sell those logs to the companies making open-weight models. The open-weight model developers then train on those logs to 'distill' a model.
- singingtoday 2mo agoFantastic, thank you for the explanation!
- ForHackernews 2mo ago...good for them? We used to call that competition. Imagine making this argument with a straight face in any other industry: "The claim is that Japanese car companies are buying Ford vehicles, and then leasing them to American consumers at cut-rate prices. In return for the cheap cars, the customers are letting the Japanese observe their driving behavior, studying how they use their F-150 and then the Japanese car companies are applying that data to design new vehicles that will directly replace Ford!"
- PersonalJarvis 2mo agohttps://www.microsoft.com/en-us/corporate-responsibility/topics/open-weight/ https://www.microsoft.com/en-us/corporate-responsibility/top...
- amazingamazing 2mo agoSadly until china scales production of hardware it really isn’t economical to run this stuff yourself. It is good it exists though to put pressure against the labs. Honestly imo this is just proof apple will win in the end. Eventually a phone will be able to run a model good enough to do most things and it then is game over.
- deleted 2mo ago[deleted]
- root-parent 2mo ago>> do most things and it then is game over. For the Hyperscalers...and Oracle...cant wait for the day...
- esseph 2mo agoAnd it floods the market with millions looking for work
- JumpCrisscross 2mo ago> it really isn’t economical to run this stuff yourself Quantised models running overnight go most of the way for non-coding tasks.
- sschueller 2mo agoModel-on-Chip is coming. GPU are for general computing but have a huge bottle neck for doing model inference. Even not being able to significantly update a model that is burned on a chip the performance gains are immense. You also don't need the latest chip fabs to make them drastically reducing the cost.
- nicce 2mo agoMany years until consumers can buy them at reasonable price. Nvdia and AMD are making GPUs bad in purpose for consumers so that nobody can build a datacenter from them. It will take a long time.
- kalu 2mo agoThe sentiment in this article is nice. But open source software is a weak analogy for frontier models. Principally because software requires zero capital investment (actually zero) while frontier models demand billions. Open models can only survive in the long run if they can (eventually) generate significant cash flows or if they are paid for by governments. Now China essentially has a monopoly on open weight models. And so supporting open source models means either supporting long term economic capture by China or supporting Chinese government control of your intelligence. Both of these outcomes are unequivocally bad from an American perspective. If you live in the valley and benefit from the US venture ecosystem you should be highly skeptical of open weight models. Banning them may very well be the best course of action.
- noncoml 2mo agoOpen Source is free as in speech Open Weights is free as in beer
- singingtoday 2mo agoAh, a true scholar!
- deleted 2mo ago[deleted]
- deleted 2mo ago[deleted]
- danny_codes 2mo agoHilariously bad take. Open weight models can be retrained of fine-tuned, that's the entire point. The idea that the "Chinese government controls your intelligence" is laughable in the case of open weight models. Once the weights are released you can do whatever you want with them. The idea that there's economic capture by the Chinese for products they're literally giving away is stupid to the point of inanity. I can only assume this account is pure shilling for the closed-source AI labs.
- 2mo ago
- cheriot 2mo agoOpen-weight and OSS are wildly different and the article makes a poor comparison. What's the incentive for the Chinese labs to continue releasing weights 5 years from now? It's not a stable equilibrium and cannot last. - The lab spending large sums on research and training does not get the inference revenue to fund those efforts. - Unlike OSS where a single volunteer can keep a project going, training costs run into the $billions. - OSS is often a two way street where features and integrations are built that the original author benefits from. Open weight models are largely a one way street because the marginal benefit is so much less than training costs. In the short term, it means Chinese labs can attract talent and, I suspect, funding from their gov. Similar to every other industry the CCP subsidized to take over.
- aleph_minus_one 2mo ago> - Unlike OSS where a single volunteer can keep a project going, training costs run into the $billions. Just some thought: Wouldn't it make sense to build some kind of volunteer computing project to train the next-generation LLM by volunteers, similar to the BOINC [1] projects or Folding@home [2]? N.B.: BOINC was particularly famous for SETI@home (completed), Einstein@Home, Rosetta@home and PrimeGrid. I still remember the time when Einstein@Home was in its heyday, and many people who loved putting together fast PCs contributed sometimes even for the reason of showing off in the statistics [3]. --- [1] https://en.wikipedia.org/wiki/Berkeley_Open_Infrastructure_for_Network_Computing https://en.wikipedia.org/wiki/Berkeley_Open_Infrastructure_f... [2] https://en.wikipedia.org/wiki/Folding@home https://en.wikipedia.org/wiki/Folding@home [3] https://einsteinathome.org/de/community/stats https://einsteinathome.org/de/community/stats
- cheriot 2mo agoI’ll be impressed if somebody can make that work considering the vastly larger compute required.
- aleph_minus_one 2mo ago> I’ll be impressed if somebody can make that work considering the vastly larger compute required. I think you underestimate the computational ressources that the mentioned (and similar-kinded) scientific projects needed. Also consider how much computational ressources people invested into cryptocurrency mining. No, I think the reasons are different: - Many companies that train AI model use training data which must not be distributed for copyright reasons (and using it is a legal gray zone)x. - Also consider that the amount of training data is insane. Scientific projects (and cryptocurrency mining, too) have the property that typically the amount of data (storage requirements) is small (or at least the computation can be partitioned so that each sub-task needs little data), but the required computing ressources are insane. - AI companies consider a huge part of their training data as their "secret sauce" (they often even paid lots of money to generate it, for example by paying world-renowned experts for writing an answer for some important question). Thus: Yes, the required computation ressources are huge, but this is a problem for which I consider it to be plausible that it can be solved. The real problems are in my opinion different.
- pianopatrick 2mo agoEventually I think to truly be like Kubernetes, you would need an AI model that has public training data and that a lot of companies collaborate on. Might make sense eventually. Same logic as companies working on Linux. "An AI model is a business necessity. But making an AI model is so expensive we should not make our own. So let's just use the open one, and contribute the stuff that we need."
- chasd00 2mo agoFTFA: American labs need to release frontier-grade open-weight models under licenses that startups can actually build on. oh now i see, the Chinese government is funding the training and release of their best models to pressure OpenAI, Anthropic, and others to do the same for competition's sake. I don't buy it, this seems more like a way to get SOTA models RL'd to comply with Chinese government approved information distribution. If I have to trust a black box of answers to questions i would trust one from a US for-profit publicly traded company subject to market forces over one approved, and heavily subsidized, by the Chinese government.
- aliasxneo 2mo agoWhat tools do we have to countermeasure the state sponsored bias in the Chinese models? Doesn’t seem like a smart plan if individuals can just compensate for the bias.
- hedora 2mo agoAlso, the choice right now is between an open weight Chinese model that is hypothetically censored to block / sabotage routine engineering flows vs a closed weight service that is definitely censored to block / sabotage those things. First anthropic guardrails blocked totally normal stuff on fable and knocked you down to opus. At this point, they kick you off fable, then opus, then sonnet. Claude then automatically builds up memories of techniques to bypass the guardrails in my long running sessions (the coordinator agent notices the subordinates got shot in the head and their sessions were pulled from context, so it parses out the lost context from ~/.claude json files, then reformulates parts of the task and uses partial results until the guardrail doesn’t trip. If I were paying for the API, this dance would cost $50-100 a pop, but I’m not, so whatever (for now).
- johnvanommen 2mo agoI worry that AI will be so fundamental to how we do things in the future, companies can mold human behavior via access to the AI tools. For instance, I worked at FICO. When I mention this, people wonder what they do. The average person only knows FICO as a “score.” FICO was founded in Silicon Valley. The average person doesn’t think about how credit scores work, fundamentally. It’s just software, at its core. FICO incentivizes certain behaviors. Ever been banned from an online forum? Now imagine if a corporation could shut you off from a technology that’s literally indispensable. Same idea.
- hsienchuc 2mo ago[flagged]
- maayank 2mo agohttps://web.archive.org/web/20260725184440/https://tobi.knaup.me/2026-07-25-open-weight-ai-is-having-its-kubernetes-moment/ https://web.archive.org/web/20260725184440/https://tobi.knau...
- Sammi 2mo agoKubernetes is a system/infrastructure orchestration tool. I completely fail to see how it is comparable to open weight neural nets. In either application or function. I'm sorry to do that hn comment thing where we all just race to contradict or talk in opposition of whatever was said before. I'm aware. But really guys, was this article really not just a miss?
- ozgung 2mo agoEveryone is talking about banning Chinese models but nobody talks how it is feasible to ban them. I think it’s impossible simply because technically there is no such thing as a “Chinese model”. There is no way to tell apart an “American” model from a “Chinese” one by looking at their weights. Weights are just numbers and you can’t assign country of origin to numbers. One can find very easy workarounds to any naive attempt to ban them by origin. So, any solution to this “problem” must include ALL open-weight models. As far as I understand this is exactly what they intend to do. Axios article linked in the post mentions that. As in this quote: “The source described leading AI labs or their allies approaching the administration every 3-5 months with an idea to ban open-source models.” It doesn’t say “Chinese” open-source models. Because they already know that it’s not feasible. Any regulation must cover all the models. Now there are solutions for that latter problem. But they are all ugly and restrictive. Making a DRM-like license protection system mandatory can be a solution. If a company wants to run an open model in their own servers, they can only use approved and certified pure “American” models. This of course creates a monopoly for the big labs who are authorized to train and distribute such “open” models. A company can fine-tune the model for its own needs but of course can’t distribute the derivative model. I’m sure there are other solutions but all of them would be equally ugly. Also these regulations can’t be enforced to other countries easily so only Americans will be restricted.
- kloop 2mo ago> “The source described leading AI labs or their allies approaching the administration every 3-5 months with an idea to ban open-source models.” That's going to hit first amendment grounds pretty quick, the same way that software in general did. The modern version of the decss flag will be a character that says "I think good weights are {...weights go here...}" They could, however, ban any payment to a chinese entity, or any entity owned by a chinese entity for inference/ai services/etc
- isityettime 2mo agoUh, how big would such a flag have to be?
- drnick1 2mo ago> American labs need to release frontier-grade open-weight models under licenses that startups can actually build on. To be fair, OpenAI has released a couple of (then very good) OSS models. I run the 20B version at home and it is excellent for reviewing text and common tasks like drafting bash scripts. There is a larger 120B that you can't realistically run on consumer hardware at reasonable tok/s too. I wish OpenAI updated these models more frequently though.
- stefan_ 2mo ago"We have been having extensive discussions around open source strategy. [..] one thing we'd like to do soon is to create a language model with the approximate capability of GPT-3 that can run locally [..]. In general, we think this helps discourage others from releasing similarly-powerful models, and makes it harder for new efforts to get funded." - Sam Altman emails OpenAI board, 2.5 years after GPT-3 Chinese release open models to drive the state of the art, OpenAI and crooked Sam do it to keep you down.
- drnick1 2mo ago> Chinese release open models to drive the state of the art Are you sure it isn't just another form of Chinese industrial policy? China does not have frontier labs, but through distillation and their own work they can get pretty close. It's not enough to be competitive with Anthropic and OpenAI, but there is still money to be made by selling compute (software as a service), and in any case it's better than being left behind in the AI race.
- potwinkle 2mo agoThere are coding tasks where Kimi K3 outperforms Fable (haven't done much comparing between it and Sol) and the Chinese labs have access to their own synthetic datasets, along with frontier research. We're leaving behind the days where Chinese models are distilled Claude, but I hope Anthropic/OpenAI can continue to accelerate.
- bigyabai 2mo ago
- Danox 2mo agoYes, and yes, again the only way to compete is to build the best not hide in a corner and once again the rest of the world will go on in AI without the United States if we flub it. Circling the wagons, isn’t the long range answer.
- __MatrixMan__ 2mo agoThis is such an obvious conclusion. To take it a bit further... Scale matters for these things. If we divide the available chips among 5 competing companies we end up with models that are trained on 1/5 of the resources that they otherwise could've been. Let the companies take turns training on shared hardware, force them to publish results in the open, and then reward them based on how well the resulting model performs at democratically chosen benchmarks. Meritocracy not monopoly. Make it about how well you wield the silicon, not how much silicon you wield, and make it a positive-sum game. If the people's data is going in, then the people should benefit from what comes out whether or not they have a subscription. If capitalism as we know it can't complete, so much the worse for capitalism as we know it.
- SilverElfin 2mo agoWhat’s interesting is no one is talking about political censorship in models and how DeepSeek, Kimi, and the rest have to abide by CCP rules. It’s a big opportunity for China to control information.
- Danox 2mo agoAnd that censorship will fail too, like sanctions, tariffs, and keeping down open source software or exchanging ideas across borders, ultimately it will fail, but that won’t stop governments across the world from trying.
- Covenant0028 2mo agoThe primary customers of these models are enterprises, and the most common use cases are office work and coding. How often do the questions of Tiananmen Square or the Uighurs become relevant in those contexts?
- applicative 2mo agoEveryone just keeps assuming, as if it were the law of gravity, that China will continue in perpetuity to deliver the weights of its 'frontier' models to Hugging Face. Its Mythos moment is a few months away and there is plenty of reporting suggesting their response will be similar, which is anyway obvious. It baffles me that anyone can seriously believe that China is going to put its Mythos successor on Hugging Face and let us all strip the guardrails etc. Things just didn't turn out the way OpenAI and Anthropic thought; nor did they turn out the way China thought.
- nunez 2mo agoThey very well might. Read up on the infrastructure projects China has provifed to Africa in exchange for oil refineries and drilling sites across the continent: https://oilprice.com/Energy/Crude-Oil/China-Is-Rapidly-Expanding-Its-Oil-Resources-In-Africa.html https://oilprice.com/Energy/Crude-Oil/China-Is-Rapidly-Expan.... Roads, schools and more.
- applicative 2mo agoThe headline from 7 years ago has not borne out; Africa is a minor supplier. The point is quite different: that everyone is harmed by advanced open weight agentic models even as we are benefited by them. This holds of African states as of any other. Many of them, by the way, are under attack from well-funded moderately insane groups like JNIM. Also by the way, groups like this are famously really good at recruiting the specific psychical type of the engineer. It is not possible that releasing open weights will continue forever, or probably even to the end of the year.
- regexorcist 2mo ago> there is plenty of reporting suggesting their response will be similar There is? Xi Jinping himself openly talked very recently about China's commitment to open weights models.
- HDBaseT 2mo agoXi Jinping realizes that open weight models accelerate open development. There is more players in China than in the US. If everyone keeps learning off each other, eventually one of the China labs will overtake the US, but AI models aren't static (anything but), and the US will leap frog in a weeks time. Therefore the constant need for open weights remains.
- pich 2mo ago[flagged]
- nunez 2mo agoI respect Knaup's opinion but disagree with the claim he makes here. Kubernetes took off because everyone could run it on pretty much anything. Bigger hardware meant bigger clusters, but developers could spin up a cluster on their laptops and deploy their apps into it. More importantly, companies could repurpose their decommissioned servers as k8s clusters, a massive unlock seeing how huge companies had heaps of these in their racks. This isn't possible with open weights models. Not in the same way. First, you're out of the game if you don't have a data center class GPU (or it's sort of prosumer equivalent). Model servers support CPU inference, but you might as well watch paint dry as you wait for results...and you'll still have to run super quantized low-parameter models that aren't as good. Realistically, companies will need to purchase millions of dollars of nVIDIA gear (through suppliers) to serve agentic-capable models at scale, an activity that is being made more expensive and complicated by the day as the hyperscaleds slurp up the demand. It also really is a huge problem that all of the open weights models are coming out of one country that also happens to be a superpower. I don't think this can be handwaved away, and it's concerning to see so many folks here minimize this. American companies running on Chinese intelligence. As a country that prides itself on being the knowledge capital of the world, the optics manifested by this are horrible, not to mention the absolutely massive supply chain risks (the counterfeit Cisco devices comes to mind, except worse because the rangers are in the weights and might not be possible to distill out). This is a threat even if you focus solely on the individual developer. Recall how AWS became...AWS. They "fanatically" focused on the developer experience. They designated this as key to their growth strategy, and rightfully so. How people talk about using Qwen or Kimi and the like on here feels like that (ignoring how these models are ALSO from huge for-profits). The only solution here is for the big labs to make some of their most capable models open-weights. There is a lot of secret sauce around routing, inference, hosting and other stuff that makes them work as well as they do, but the community can figure that out. Of course, this is basically a death sentence to those companies, but that's what they get for playing with monkey paws, I guess.
- samizdis 2mo ago> ... once an open platform that people can customize becomes the industry’s center of gravity, no single vendor can match the combined rate of innovation around it That assertion is sublime, and if not true, it should be and can be. Thanks for the thought (and the optimism hit).
- nedt 2mo ago> Kubernetes did not win simply because its repository was public. And here I am still using docker with compose because it's easier and works. Always saying I might be looking at swarm when needed, but then never really need it. Just because you are using it doesn't mean everyone is and doesn't mean it has won. On the other hand yeah I want open models and I want them local as well.
- shard972 2mo ago[dead]
- bmiekre 2mo agoThis seems like a pretty balanced and thoughtful idea. I’m HN will hate it
- dboreham 2mo agoIt's going to be so complicated that nobody quite understands how it works?
- cmdrk 2mo agoDrat, I was hoping this article would be about outlawing Kubernetes.
- Alien1Being 2mo agoMeanwhile America is having its glorious McCarthy moment....
- xivzgrev 2mo agoI worry about the Chinese models sending data back to China. How do we know that isn't the case?
- ehnto 2mo agoDepends on how you're running them I suppose. If you are running them on local hardware, then the harness is what you are concerned about. A compromised local model might be able to exfiltrate data via tool calls. If you are running it in a cloud, the above is true, but also you have to trust the cloud provider isn't relaying your chats to a third party, for training or perusing. If I am totally blunt, if privacy is a concern then corporate US has proven time and time again to be a terrible steward of your private data. You should assume it's being ingested or sold, or turned into metadata that's sold in a data laundering pipeline. If it's corporate secrets, on-prem and enterprise offerings are your options. Enterprise at least puts the providers on the hook legally, when they inevitably fuck up keeping your data private in some way.
- Retro_Dev 2mo agoIndeed - if you need to be absolutely confident about security and privacy, run a model locally and audit the inference software and potential tooling.
- deleted 2mo ago[deleted]
- zackwu 2mo agoThere are zero-retention 3rd party providers (EU/US companies), but it's your choice whether to trust their claims or not.
- HDBaseT 2mo agoI worry about the US models sending my data back to Israel.
- cindyllm 2mo ago
- jitbit 2mo agoYou can contribute to an OSS platform. To accellerate innovation, adoption etc You CANNOT contribute anything to a model.
- memedesimo 2mo agoIt is funny to see how you feverishly seeking "legitimate ways" to suppress successful concurents. I would call you "sore losers", but as I know how insane you actually are, you might just start a war. (Yes yes, they "stole from you", and "Tiananmen square", and "Uigur genocide", and.. and I just can't wait for the monent someone locks your crazyness behind the bars of an asylum.)