9 ms·
The gap between open weights LLMs and closed source LLMs
- samat 3mo agoArticle confuses open source models with open weights models. Not the same thing. It’s used right in the articles body, but title is misleading.
- NitpickLawyer 3mo agoLiterally no one cares. There are "full" open certified GMO free grass fed training data blah blah models. Apertus, Olmo, etc. No one cares. For all intents and purposes people use the term to describe a model that you can run locally and are allowed to modify and re-release. The rest is useless semantics. No one can "rEpRoDuCe" a model anyway.
- throwuxiytayq 3mo agoNo-one cares to quit social media or stop using Windows, but it’s a goal worthy of discussion all the same. The name is bad, doesn’t even make any fucking sense and it gives open source a bad rep.
- komadori 3mo agoI wouldn't say that no one cares, but obviously many fewer people care when the cost of "recompiling" a model from its open source training pipeline is so high. Also, if you only have the weights, you can still use it to generate training data for a new model (i.e. distillation) so it's inherently less locked down then closed source binaries were.
- judge2020 3mo agoopen source vs source-available. Companies taking an extremely cautious approach to AI can't use source data that is potentially a violation of copyright (pending worldwide court decisions and/or regulation on said topic). Although that cat is already out of the bag for basically every stock-traded company using LLMs trained on non-licensed data, so I don't see there being much actual risk in using them.
- reinitctxoffset 3mo agoI was advocating for "available weight" as a value neutral term for a while. I gave up. No one cares. And no one will ever tell the truth about the training anyways. Substantial and growing freedom beats zero freedom ever again.
- jackconsidine 3mo agoAchilles and the tortoise [0] is usually a fallacy. If the tortoise has a head start, then Achilles will never catch it because in the time it takes Achilles to reach the tortoise's location the tortoise has moved some degree further, ad infinitum. Obviously not real because Achilles will pass the tortoise -- I think a fallacy because the framing creates a fake asymptote (they will both pass the point where they're approaching a tie). In this case it may actually apply though, no? Open models get better from closed model distillation? [0] https://en.wikipedia.org/wiki/Zeno%27s_paradoxes https://en.wikipedia.org/wiki/Zeno%27s_paradoxes
- igravious 3mo agocomparing a thought experiment about relative movement through an alleged continuum over the sum of infinitesimal quasi-instants to the release cadence and maturation of open weights to proprietary LLMs is super bizarro guy
- profsummergig 3mo agoIMHO, the biggest problem with the future of open weights models is that currently, open weights models are the result of philanthropy by some private org. (e.g. DeepSeek). The spigot can be turned off at any time. Until there's some sort of "community owned hardware", open weights models are always at risk of being discontinued.
- Shitty-kitty 3mo agoIt's just a smart business decision that allows their models to compete and gain market-share against much pricier private models. No philanthropy there.
- foxglacier 3mo agoIt depends how you define philanthropy - obviously corporations don't just donate such valuable products to the world to make it a better place, but in effect that's what they end up doing in their effort to gain market share or brand recognition. Actual human philanthropists are sometimes doing it for the similar reasons of self-promotion.
- Shitty-kitty 3mo agoOpen source, Open weights, these are core business decisions.
- NitpickLawyer 3mo agoYeah, but the biggest plus for open models is that they can never be taken away. In other words, whatever capabilities they reach (even if there will never be another model), those stay forever. That can't be said for API-based models where a provider can sunset models whenever they feel like (i.e. gpt5-mini will soon be gone, and replaced by a more expensive 5.4-mini, same for goog, etc). And there will always be incentivised parties that release models. Nvda for one has every incentive to keep the nemotron line going, as they're directly profiting from people running this. And the models aren't really far from open SotA anyway. Goog will probably continue to release the small models, since they'll use them for browser stuff anyway, and know that they'll leak. So for them it's a win-win to release the small models and gain some dev market share. And the chinese labs also have incentives to keep releasing models, and will likely continue to get gov support to do so (yay commercial wars between nations).
- jacobgold 3mo agoIt would be interesting to know how much of a boost the closed models companies are giving the open models. If the closed models stop improving will the progress of open models slow?
- amluto 3mo ago> It would be interesting to know how much of the "distillation" boost is helping the open weight models keep up. Some people in China surely know. > Like if the closed models stop improving will all the closed models also stop improving? Seems extremely unlikely, unless the models all hit some kind of wall soon. The Chinese companies may be behind the US in compute capacity, but they have excellent researchers [0] who are probably approximately as good as their US counterparts at the kind of problem generation and RL that is currently working so well. I would be very surprised, though, if the models cannot continue to be improved rapidly in any area that allows a tight feedback loop like programming, at least up to the point where we puny humans lose the ability to define objective functions. (And, conversely, I don’t expect magic in fields where the feedback is slow or expensive. A model is not about to reliably invent a wonderful medicine for the same reason that a large and extremely competent pharma company cannot: the evaluation process is extremely slow and it’s so expensive that the kind of utterly enormous corpus that is driving the current progress in coding is simply not available. Running RL on m iterations of n medication-development trajectories each is going to cost n*m times $10-100 million and take m years if it’s even possible at all.) [0] The US advantage in this space will likely decline, since the brain drain from the rest of the world via the US university system to US labs is drying up.
- typs 3mo agoPerhaps. RL env companies based in the U.S. sell to Chinese labs quite a bit too though (though on a discount, once they're no longer on the frontier)! And it would make sense that a lot of these problems which are based on work in the U.S. enterprise economy would be coming from the U.S.
- amunozo 3mo agoWhy are we assuming only American labs can innovate? DeepSeek already innovated a lot in efficiency, for example.
- justindotdev 3mo agoat first glance, these graphs are confusing
- nsingh2 3mo agoYea these plots are too noisy and dense. Especially that second one, lines all over the place.
- gunalx 3mo agoUtterly unreadable on mobile
- taffydavid 3mo agoTotally agree. That first chart is just four vertical lines, I can't figure out what the hell I'm supposed to be looking at without reading the article.
- gehsty 3mo agoInteresting to consider this inline with recent us export bans, could the US be squandering its lead by giving the open source, largely Chinese labs catch up (in terms of model quality available to masses), will US labs be able to maintain the lead without users being able to use their latest models?
- ggm 3mo agoWhy do you think this matters? Not that it does or doesn't but what quality does "US WINS" or "CHINA WINS" bring to the table?
- knowaveragejoe 3mo agoI think the unspoken fear is that if we assume one or the other will "win" in reaching AGI(or whatever threshold of capability), the rest of the world will sooner or later live under their system of rule as a consequence
- ggm 3mo agoI very much doubt the primary reason nation states are lining up to permit or forbid access to these systems is 'fear of future AGI dominance' I think it's much more immediate/present: the weights and the information breach significant strategic controls on national data and posture, which can be back-derived from the models. If you can analyse a model, you can infer what structural inputs dictate it.
- knowaveragejoe 3mo agoCan you expand on that reasoning more? Claude etc has national secrets hodden somewhere in its weights?
- gehsty 3mo agoI think the issue matters (not us vs china) due to the investment and exposure normal people have to the valuations of the AI companies. It feels like the US govt could make this the pin that pops the bubble. If these companies loose their lead their value drops and the stock market tanks.
- llmslave 3mo agoThe gap is huge and im tired of reading these articles constantly
- Gigachad 3mo agoAre you talking about hosted vs the ones you can easily run locally? Because there are open models that require hundreds of gb of vram which are apparently pretty close.
- verdverm 3mo agoon the Will It Mythos benchmark, small models are punching way above their weight(s) gemma4-26B (#7) qwen-3.6-27B (#9) https://news.ycombinator.com/item?id=48640196 https://news.ycombinator.com/item?id=48640196
- Gigachad 3mo agoI've tried running qwen 3.6 locally and it felt like LLMs a year ago where you can get them to do some stuff but the tasks have to be very small and you have to course correct them a lot to the point it's hard to say it's any faster than doing it all yourself. Certainly the gap is closing but I feel it still makes more sense to pay pennies to run the full sized open models hosted on much better hardware.
- verdverm 3mo agoI had qwen36moe revamp my PhD thesis with a rewrite using JAX. Gave it access to my old code, helpednitnwhen it got stuck or didn't quite understand a few times. Overall I was very impressed with its open box reimplementation. I remain of the mind they are widely underrated.
- trollbridge 3mo agoWhat edition of Qwen 3.6? 35b-q6 (with MOE) has felt good enough for general purpose agentic coding. You obviously need a 32GB card to run that (or a 64GB Mac), and realistically anything lesser than an RTX 5090 is going to be too slow for practical use.
- JumpCrisscross 3mo agoNow let’s look at the economics of buying versus renting. I’ve seen a lot of attention given to hardware capital costs. But a comment the other day got me thinking about power costs, too—at what performance differential do these factors intersect to make on-prem economically competitive with datacenters for businesses?
- deleted 3mo ago[deleted]
- dabinat 3mo agoI believe the open model party will eventually end. Perhaps because companies realize it’s too much of a commercial advantage, countries don’t want to give other countries commercial or military help, or maybe even an outright ban after someone uses an open model to guide them through how to make a bomb.
- stkdump 3mo agoPossible. Though I think open source innovation of hobbyists is currently hampered, because of open weights releases by chinese labs. I believe once the labs stop doing this, a globally coordinated open source ecosystem will fill the gap.
- taffydavid 3mo agoIf we were going to ban technology because it helped people make bombs we wouldn't have access to much anymore.
- dabinat 3mo agoSure, but this administration doesn’t behave logically and big AI companies are already pushing the “AI is dangerous and only we should be trusted to wield it” angle.
- _pdp_ 3mo agoFrankly it does not matter if there is gap because for most practical use-cases the end user can barely perceive the difference in intelligence. On paper frontier models will be ahead of the curve but I don't think hardly anyone will be able to tell if a piece of work, say a landing page, is created with Fable or GLM and that is the point. The perceptible intelligence will reach a point beyond which it is no longer considered, except for some narrow use-case.
- nomel 3mo ago> except for some narrow use-case. I think it's entirely the opposite. For narrow use cases, like web pages and crud/GUI, the open source models don't show much of a difference.
- mft_ 3mo ago100% agree. My impression is that the open-weight models have been drawing close-to-level at coding tasks, while Anthropic and OpenAI have been putting large amounts of effort into developing their models' abilities in other domains: legal, biomedical/science, etc. Anthropic (especially?) has also been putting more obvious resource behind optimising their harnesses - from Code to Cowork (which is kinda Code for normies), Design, etc.
- _pdp_ 3mo agoGLM 5.2 has replaced "normie" agentic workflows previously backed by Sonnet and Opus. So I don't know. From my end it seems to me they are perfectly capable of working agenticly.
- mft_ 3mo agoMaybe we have different definitions of 'normie'. I'm talking about people who aren't in IT, and who are maybe just learning to use LLMs for aspects of their daily work. These people only know of the big three models, at best - they very rarely know of the open-weight models, and would even more rarely (given their model access is likely determined at a corporate level) be able to access them.
- christina97 3mo agoThe Chinese models will not overtake the frontier US ones given the current way things are going. The US models derive their lead from incredible efforts to source more and higher quality (mostly synthetic data) via great feats (eg generating with humongous teacher models that could never feasibly serve interactive traffic). The Chinese models advance via heroic efforts to optimize models and great feats to secure more and higher quality training data from the US frontier models. For an (Chinese) open weight model to surpass the (US lab) frontier models, this equation must flip and the Chinese labs must entirely retool from harvesting frontier model data to producing the data systems and efforts to produce novel data; as well as procuring latest generation hardware en masse for this. This does not happen easily. Also training a frontier scale model is actually not such an unimaginable feat: doing all the inference with the teacher models is where the hardware goes.
- andy99 3mo ago> Chinese labs must entirely retool from harvesting frontier model data to producing the data systems and efforts to produce novel data Even if your characterization is accurate, they could do this tomorrow and are not so myopic that they wouldn’t have thought about it. I don’t see this as a barrier, and I see a lot of the same underestimation of Asia that’s been happening for 50 years. There’s not some innate American advantage to building LLMs, and personally I think whatever head start the US has is going to be squandered on delays from the export control “to dangerous for release” LARPing we’re seeing.
- s1artibartfast 3mo agoWhy would those have any impact on R&D speed? Most are funded and close to cash flow positive
- ant-kinesthetic 3mo agoExactly. If they wanted to they could produce the same amount of data. Companies like Scale, Mercor, Surge exists for a reason, a reason that doesn't need to exist in China if they mandate Chinese enterprises to provide all their real world data (or have them work inside RL environments) to the model companies for post training. There is no real advantage that US companies have except a head start, and as Jensen said, a ton of the research advantage is skewed since a lot of the best researchers in the US are Chinese nationals. I do think the model is just one piece of the pie (not to echo Jensen too much), and hopefully we will always be able to serve these bigger frontier models in a much more efficient way as well as building out the application layer faster which actually makes them useful and/or more dangerous/powerful.
- doctoboggan 3mo agoIf the Chinese government is as involved in LLM development strategy as many people claim, wouldn't you expect them to immediately cease releasing open weight models and restrict access as soon as they start producing the frontier models? I am assuming this is what the USG thinks and is why they are trying to cut off the flow to foreign nationals ASAP. LLMs are an undeniably valuable tool, and governments like to control those.
- sdesol 3mo agoI talked about this before but China would be in much better position if LLMs turn into a commodidty. Where they can dominate is in hardware, as fast and cheap inference is probably going to be the moat.
- verdverm 3mo agoMy futurology is that most of us will end up on unlimited token plans like we are for mobile data. We don't need the very best model for most tasks and the trend in computing has always been towards cheaper and more efficient unit economics. I do not see this ending any time soon.
- nicce 3mo agoHow do you know that Chinese don’t have powerful private models already? Maybe they just allow opening the ”bad” models…
- psychoslave 3mo agoAs far as we don't know, they might also have operational stargates large enough to let their starships pass through. Actually, every country might have that. But what is impactless to the wider world will always be as significant as something that never existed.
- eunos 3mo agoXi Jinping isnt as AGI pilled as US govt. CapEx in US is significantly focused on AI related things like chips and data center. It's more diversified in China as they also invests hugely on renewables, EV, BESS, etc.
- maxiniol 3mo agoAm I the only one flagging inconsistencies in the different evaluations on the 18 benchmarks ? Why is sometimes the closed frontier model grok ? And then opus 4.8 ? Compared to GLM 5.2 once or sometimes Kimi 2.6 ?
- tzs 3mo agoI wonder if a lot of the companies and governments that seem to think it is essential to be on the forefront of applying leading edge LLMs to the point of starting to become dependent on them are going to find themselves in a situation like that from the Arthur C. Clarke short story "Superiority"? [1] [2]. [1] The story: https://nob.cs.ucdavis.edu/classes/ecs153-2019-04/readings/superiority.pdf https://nob.cs.ucdavis.edu/classes/ecs153-2019-04/readings/s... [2] Wikipedia: https://en.wikipedia.org/wiki/Superiority_(short_story) https://en.wikipedia.org/wiki/Superiority_(short_story)
- cedws 3mo agoI haven’t seen it discussed anywhere that closed models can essentially cheat benchmarks right? What Anthropic or OpenAI brand as a model doesn’t necessarily have to be just weights, it can be a whole backend system that augments the model itself. With this they can score better benchmarks than an open source model that is weights alone.
- snthpy 3mo agoGood point
- jstanley 3mo agoSure, I think that's fine, that all counts. It counts for open source too, it's not like they're somehow running these benchmarks without any harness. Nobody cares if your AGI is 100% made out of neural networks or if it's like 50% neural networks and 50% perl scripts.
- stkdump 3mo agoI think they mean cheat in a Dieselgate sense. You detect that you are being tested with a specific benchmark question and heuristically give the correct (manually programmed) answer. That wouldn't be AGI.
- ChrisArchitect 3mo agoRelated: The unbearable cheapness of open weight models https://news.ycombinator.com/item?id=48668255 https://news.ycombinator.com/item?id=48668255
- casey2 3mo agoThis is just and example of "lying with statistics". Going by compute efficiency the gap has already closed (both in training and inference coincidentally).
- linzhangrun 3mo agoUSA, a country that known for the land of freedom, is now restricting frontier models to the point where non-Americans cannot even use them. China, a "authoritarian state" country, "the antonym of freedom", with a software industry that is especially capitalist, has produced all the competitive open-weight models. It really is IRONIC. Disclosure: I am Chinese, and I understand this strategy comes from being behind, using open source as an asymmetric way to compete and make up for missing compute by sharing the burden, etc. But still, very ironically.
- mft_ 3mo agoYour comparison falls apart in the first few words: > USA, a country that known for the land of freedom The US might say it's the land of freedom, but it's been playing the game of economic protectionism for centuries. This is just the latest example.
- sinatra 3mo ago> it's been playing the game of economic protectionism for centuries. Equal to or more than other major countries? If so, can you show me anything supporting your claim?
- mft_ 3mo agoI wasn't making a comparison or a value judgement. We're obviously both aware that other countries (and blocs, like the EU) also play the same game. Rather, I was simply noting that the US' "freedom" branding that the poster was referring to doesn't necessarily apply in the way they were assuming (their point was that the "free" US restricting models more than "authoritarian" China was ironic) given the US' history of economic protectionism.
- zb3 3mo agoI just hope CCP doesn't follow the US government and won't pull the plug before their companies release something on-par with the US frontier models. The question is whether US models not available to the general public will count. The question is not whether they'll prohibit open-weight models better than the US ones, because we all know the obvious answer.
- StreamCtx 3mo ago[flagged]
- mft_ 3mo agoIf the belief that open-weight/Chinese models depend significantly on distillation of the latest frontier models is correct, then presumably the gap will stabilise to the minimum time required for extraction of meaningful data (from the latest frontier model) plus finalisation of training of the latest dependent model. This gap can be minimised by increasing the process efficiency, but can't be eliminated entirely. (Attempts to hinder distillation from Anthropic/OpenAI may shift the balance too.)
- sinuhe69 3mo agoAt this point, I think open weights vs proprietary models is a misnomer. First, we can not be sure the next release will remain open weights as Qwen 3.7 has showed. And second, they are all Chinese models. So instead of open weights, perhaps Chinese AI models is a better word choice.
- trollbridge 3mo agoIt's not China's fault they're the only country releasing open weights. Qwen has always alternated having an open release followed by a "max" release that isn't open weights.
- sinuhe69 3mo agoAlthough I can see that some people perceive “China’s model“ as a bad connotation, it’s simply the truth. So face it.
- taffydavid 3mo ago> Now is probably a good time to liquidate your pension, fly to a remote island somewhere, and live out the remaining 6 months or so of civilization in peace. > So maybe the open source apocalypse won’t happen yet. Sorry I wasn't at the last doomer meeting, when did we decide good open source models are a harbinger for the apocalypse?
- kageroumado 3mo agoIf anything open-source models are a hedge against the apocalypse. Or at least against the cyberpunk dystopia.
- moffkalast 3mo agoIt's an apocalypse for the paid SAAS providers, which is the best thing that could possibly happen to help prevent the completely feudalist future they want to enact down the line. The world's knowledge on a USB stick is the the greatest danger to those that would like to restrict and sell what they freely benefited from themselves.
- port11 3mo agoCute. Climate change’s apocalyptical impact on food crops and cancer rates (post-ozone collapse) never convinced people to enact change. But hey, it’s open-model LLMs, the boogeyman! Can’t have that, it must be OpenAI or Anthropic safely controlling the market and calling all the shots.
- roenxi 3mo agoAre you sure you haven't gotten your catastrophes crossed? Ozone depletion was a different crisis and people did enact change, the ozone hole has been closing fairly steadily. Wikipedia [0] thinks the prospects for the ozone layer are pretty good. [0] https://en.wikipedia.org/wiki/Ozone_depletion#Prospects_of_ozone_depletion https://en.wikipedia.org/wiki/Ozone_depletion#Prospects_of_o...
- 3mo ago
- zkmon 3mo agoWhat matters might not be gap itself. For the bulk of AI users, it's the sufficiency of the capabilities of a model, is all that matters. If an open-weight model meets their requirements and far cheaper than closed weights model, then they have no reason not to go for the open-weight model.
- kuchta 3mo agoAre there really some Open Source models there? Open Weights yes, but Open Source requires to open source of all the training data too, otherwise you can't reproduce the weights in the same manner as you would reproduce binaries from Open Source code.
- swiftcoder 3mo ago> What is notable is that a large amount of the total improvement of models has been in the coding benchmark. The coding index has gone from 15 months behind to only a month or two behind This makes sense, right? Coding is one of the most obvious short-term uses of models, it also has a readymade market willing to pay a lot for tokens, it has a huge corpus to work with, and a significant degree of validation is built into the problem domain...
- jessinra98 3mo agoCurious what other people's tipping point has been for picking one over the other