7 ms·
IMHO, the biggest problem with the future of open weights models is that currently, open weights models are the result of philanthropy by some private org. (e.g
by profsummergig 4mo ago
IMHO, the biggest problem with the future of open weights models is that currently, open weights models are the result of philanthropy by some private org. (e.g. DeepSeek).
The spigot can be turned off at any time.
Until there's some sort of "community owned hardware", open weights models are always at risk of being discontinued.
- Shitty-kitty 4mo agoIt's just a smart business decision that allows their models to compete and gain market-share against much pricier private models. No philanthropy there.
- foxglacier 4mo agoIt depends how you define philanthropy - obviously corporations don't just donate such valuable products to the world to make it a better place, but in effect that's what they end up doing in their effort to gain market share or brand recognition. Actual human philanthropists are sometimes doing it for the similar reasons of self-promotion.
- Shitty-kitty 4mo agoOpen source, Open weights, these are core business decisions.
- NitpickLawyer 4mo agoYeah, but the biggest plus for open models is that they can never be taken away. In other words, whatever capabilities they reach (even if there will never be another model), those stay forever. That can't be said for API-based models where a provider can sunset models whenever they feel like (i.e. gpt5-mini will soon be gone, and replaced by a more expensive 5.4-mini, same for goog, etc). And there will always be incentivised parties that release models. Nvda for one has every incentive to keep the nemotron line going, as they're directly profiting from people running this. And the models aren't really far from open SotA anyway. Goog will probably continue to release the small models, since they'll use them for browser stuff anyway, and know that they'll leak. So for them it's a win-win to release the small models and gain some dev market share. And the chinese labs also have incentives to keep releasing models, and will likely continue to get gov support to do so (yay commercial wars between nations).
- felooboolooomba 4mo ago> they can never be taken away Your right to 3d print whatever you want is about to be taken away (in California). What software you can run on your computer can already be restricted. Absolutely everything can be taken away. The simplest way to remove open models is probably to declare them a tool that terrorists could use. Crazy? Yes, the world is totally crazy these days.
- redox99 4mo agoThat only affects people in California. Whereas Fable being shut down affects people all over the world.
- anticorporate 4mo agoThere's also, importantly, a distinction between what are told we can no longer use, and what can actually be taken away. Open source and open hardware can be called illegal by a government, but, if we collectively invest our energy into open alternatives, they can't be taken away in the same sense. I can build a RepRap printer and I can use a local AI model. It's on all of us to make sure that the open alternatives are viable, maybe in the current global political reality now more than ever. Making something illegal isn't a disincentive for everyone. When they start banning books, some of us start assembling printing presses.
- echoangle 4mo agoBelieve me, if the government wants to stop you from having access to something like that, they could do it. Just give people some incentive to report you and make really harsh punishments and everyone will be thinking really hard about how bad they want have access.
- dvngnt_ 4mo agoThey can stop piracy or child predators. what makes you think they can prevent access to running models that require no internet access to run
- fridder 4mo agoWe need a SETI@Home but for model training
- kamranjon 4mo agoHave been thinking about this a lot lately.
- 0x3f 4mo agoConsumer hardware over the internet is not really suitable for this, AFAIK.
- baby_souffle 4mo agoThere's some really early days work on making training loops robust to failure but they all have trade-offs right now. I remain hopeful that we'll be able to democratize the entire tech stack for this tech.
- Azantys 4mo agoI think model training is pretty hard to do efficiently on a vastly distributed network. If the model cant fit into the VRAM of the node your performance becomes so bad its useless, so a distributed model could only be properly trained if the size of the model doesnt exceed the majority of the nodes VRAM sizes. Maybe there is a different way of doing training but this would be the only way I can see. And it would still be much worse than just using a big datacenter where everything is fully interconnected. BOINC projects work great because its usually just a lot of small compute and memory required so every old desktop and laptop can contribute. Training a model which can compete and is not tiny requires neither low compute or low memory amount. BOINC tasks take minutes usually or sometimes hours but not weeks or months like training a model from scratch. But something like 7B or lower could maybe be trained like this. Im not sure but I think someone is already working on something like this but I dont remember the name of the project.
- wuschel 4mo agoMy understanding is that in addition to your comment and the development of a method to separate the training data for distributed learning, the latency/bandwidth of systems connected on the internet is a challenge, too. Information has to be sent around before and after any hypothetical number crunching.
- recursive 4mo agoThis seems backwards. Access to Fable can be removed. I don't see how an open weight model can ever be put back into the bag though.
- Smaug123 4mo agoThe model itself, sure; the comment is about the production of more advanced models (to keep open weights near the frontier).
- recursive 4mo agoThe proprietary spigots can be turned off at any time also. To me, that seems more likely.
- jfaat 4mo agoI'd call 'more likely' an extremely safe take given that it's exactly what's happening right now
- ItsMonkk 4mo agoGovernments can always offer prizes, and any model that sufficiently meets the criteria of the prize would win it and claim a large cash prize. Once claimed, the model would then be free to all forever.
- ForHackernews 4mo agoIt's not pure philanthropy: https://gwern.net/complement https://gwern.net/complement
- notnullorvoid 4mo ago> Until there's some sort of "community owned hardware" Or until some bright people figure out drastically more efficient means of training.
- UncleOxidant 4mo ago> The spigot can be turned off at any time. True. And it's possible that this has already happened at Alibaba Qwen - at least for the smaller models that people had a chance of running at home (122B and smaller).
- gunalx 4mo agoWe'll see. The qwen team has always released a few close to sota but proprietary models in between tgeir open releases. We did get 3.6 35B and 27B so its not all set in stone yet. Its higley unlikely we get another open llama model though after the llama4 flop, even if their muse spark seems pretty good.
- trollbridge 4mo agoHas it though? They've been releasing free models interpersedwith the "Max" models for quite some time.
- deleted 4mo ago[deleted]
- slashdave 4mo agoTraining these models is not a "hardware" problem.
- deleted 4mo ago[deleted]
- nomel 4mo agoI think that simplifies it a bit. You can't train without hardware, which is why the Chinese companies are illegally importing Nvidia cards [1]. [1] https://www.theinformation.com/articles/deepseek-using-banned-nvidia-chips-race-build-next-model https://www.theinformation.com/articles/deepseek-using-banne...
- adrian_b 4mo agoThe usefulness of the smuggled NVIDIA GPUs has greatly diminished for AI purposes, because the elimination of NVIDIA as a competitor has allowed the growth of the production of domestic GPUs. Moreover, China has just demonstrated a supercomputer faster than any US supercomputer, which unlike the US supercomputers, which need GPUs, achieves its high computational throughput with custom CPUs designed in China (implementing an Armv9-A ISA with SME, i.e. the scalable matrix extension, and with BF16/INT8 operations for AI). The CPUs used in that supercomputer can reach both a computational throughput and a memory bandwidth sufficiently high for training any LLMs (they have fast HBM memory). Their only disadvantage in comparison with the best NVIDIA GPUs is a slightly lower energy efficiency, but China has abundant cheap energy so this is not a serious disadvantage for them.
- menaerus 4mo agoSIMD programmers have to be paid very well then in the China ... Jokes aside, some 2 or 3 years ago I thought that it is becoming inevitable for CPU designs to become an extended versions of their already quite capable vectorized execution engine units.
- deleted 3mo ago[deleted]
- jmyeet 4mo agoHow is this a complaint? Once you have the model, you have the model. Download DeepSeek-R1 671B and you have it. You might not get improvements in the future, just like you may not ever get a future release of an open source project. Is that an indictment of open source? But consider the alternative. OpenAI and Anthropic can shut off your account or API key at any time for any reason. How is this better? You have way more security when you're running your own model.
- girvo 4mo ago> Download DeepSeek-R1 671B Dunno why you'd want to though, considering v4 Pro (and even Flash) outpace it drastically
- throwawayffffas 4mo agoI don't think that's the case, it's not philanthropy, they are getting something out of it. The labs are learning from one another from the shared models. Plus I am certain it makes financial sense. I am guessing here but fully utilizing a subscriptions limits probably costs the operator more money than the subscription revenue, that is why anthropic is making such a big stink about the chinese data harvesting. By releasing the weights, you are relieving yourself from that burden because the competition does not need to hammer your subscription service they can just download your model and analyze it and run it all day. Also for the largest models it makes no sense to run it yourself unless you are a major player. Renting the hardware is ludicrously more expensive than their subscription tens of thousands of dollars. And buying the hardware to run them is in the hundreds of thousands of dollars.
- yorwba 4mo agoThe primary benefit of releasing weights is the attention it generates. Some people have the hardware to run it, try it out because it's free, tell everyone about it, and then even people who don't have the hardware might get interested and pay the original developer. So it's a marketing expense, basically. The most popular LLM product in China is Bytedance's Doubao. You probably haven't heard of them since they never released weights and don't benchmark particularly well, but Bytedance already had enough users on its other apps that they could directly advertise Doubao to.
- bijowo1676 4mo agoI believe we are still very very early in AI development, so it doesnt even make sense to close models. Open source and open weights model is how you can harness the potential of all humans to continue development and improving the SOTA of your model. Literally every student on the planet wants to play and improve these models for their own use case. Plus the ecosystem, once you have users in the ecosystem on your open weight model, this is a giant leverage point in itself
- FooBarWidget 4mo agoThat's not meaningfully different from philanthropy. If Chinese AI products generate sufficient revenue with cheaper marketing strategies, then the incentives for releasing open models will go away. Right now, there is a shortage of talented researchers, and the attention that open models generate allow them to attract good hires. But this is a fragile dynamic that can break in the future. It's not very different from commercial open source work, except it's much more capital intensive and lower volume.
- gwerbin 4mo agoIsn't this also true of a lot of FOSS software and libraries? tensorflow and pytorch for example, among many others.
- 40four 4mo agoWe should address the elephant in the room. The problem with the future of open weight models is not they are created as a result of philanthropy by some private org. All of the top contenders are created by the Chinese government. I don’t think we should describe these companies as simply releasing these highly capable open weight models out of the goodness of their hearts
- psychoslave 4mo agoBhutan didn't release any model yet as far as I know, if the level of care government give to people actual happiness is what are supposed to be concerned about here. Among over countries that are consistent being on top on gross national happiness are Finland, Denmark, Iceland, Switzerland, and the Netherlands. Among them the current abilities to release open models is observable. USA unfortunately continues to fall down quickly in World Happiness Report rank, and that's not because many other countries made great progresses.
- nmfisher 4mo agoNone of those companies are created by the Chinese government. They're obviously subject to the Chinese government, whose whims may change at any given moment, but as we're seeing at the moment, so are the American companies. And while I don't have a very positive view of the Chinese government, last I checked, they haven't been dropping bombs on innocent schoolchildren recently.
- cheesecakegood 4mo agoBombs are a bit of a non sequitur here. The point is that Chinese companies are demonstrably hostile to American ones historically (and threatening in some specific structural ways to the American consumer). The presentation may be similar but to attribute American ethics to a Chinese decision is dubious.
- defrost 4mo agoIsn't the nature of capitalism such that many companies are demonstrably competitive (aka 'hostile' ?) with one another? Chinese companies have also demonstrably pandered to the American consumer for many decades now. To further muddy the waters, US companies have, some would argue, been openly hostile to the American consumer via monopoly practices, restricting access to purchased devices, etc.
- alfiedotwtf 4mo agoExactly my worry. I’m optimistic in the future the EU, the EFF, the GNU, or the Linux Foundation could have been the umbrella to run a LARGE open model for everyone. It’s sad to think that Mozilla spent years and millions doing virtual reality and AI, they would have been perfect to do this but let’s face it - who knows if Mozilla will be around even 5 years from now
- Eridrus 4mo agoI think the bigger issue is the ever increasing capital requirements, which may cause even the closed weight companies to fall away from the frontier, e.g. Google & Meta are barely hanging on. For Google it feels a bit existential to remain at the frontier, but even then they're barely there. I hope that we find ways of continuing to improve these models besides continuing to exponentially increase capex spend until all but one of your competitors falls away.
- Onavo 4mo agoGoogle and Meta's failures are more due to mismanagement no?
- disgruntledphd2 4mo agoAt times of rapid change, having a working business model can be a disadvantage. For instance, Facebook were able to optimize their core ads product for mobile, in a way that was much more difficult for Google.
- Eridrus 4mo agoI have no idea what's going wrong inside Google/Meta, they certainly have capital. But when you need this much cash, not many people are going to be able to have a shot on goal. It would not surprise me if Meta threw in the towel. Microsoft and Apple aren't even trying.
- ehsankia 4mo agoIsn't another issue that most successful open models are distilled from closed models, but closed models are putting more and better safeguards against distillation?
- alecco 4mo ago> Until there's some sort of "community owned hardware" The hardware is already available for renting at reasonable prices. We need community funding. I wish people pooled a fraction of the money they burn on local GPU rigs on funding training/testing/etc. A big problem is like in open source: it's way too atomized. Just one competitive ground-up community LLM would require tens of millions $. But who gets to pick? IMHO the only chance is highly specialized and smaller LLMs instead. And this is still millions to train. And remember LLMs are competitive for only a handful months.
- matheusmoreira 4mo agoI wish we had some kind of distributed training capability... Like Folding@home, but for LLMs.
- woctordho 4mo agoSee the recent advance of DiLoCo at Nous Research and Prime Intellect.
- matheusmoreira 4mo agoReally interesting!! This gives me hope!
- c0rruptbytes 4mo agoDeepseek isn't philanthropy, it's a hedgefund trying to short the western AI market by saying "hey we can do 90% of they can (arguably better at a density metric) for a 1/10th of the cost" it's my theory at least, the Hindenburg Research of AI
- dboreham 4mo agoWidely held belief in investor circles is that the Chinese government has a goal to deflate the US AI bubble and Deepseek is part of the plan to achieve that goal.
- aurareturn 4mo agoWhy would China care about deflating the US AI bubble? Why do we think there is a bubble for sure in 2025/2026? Why doesn't China also worry about their own AI bubble inside the country?
- tarpitt 4mo ago>Why would China care about deflating the US AI bubble? To weaken the stature of the USA on the global stage relative to themselves. Perhaps decrease US investment in AI and slow creation of some general AI superweapon I suppose. Because the goal is to show that cheap chinese AI can compete with expensive USA AI, it's nessisarially a low-cost attack relative to the "damage" it could create. >Why do we think there is a bubble for sure in 2025/2026? Well that's the position that these chinese firms are trying to convince us of, and they can convince us by undercutting proprietary models in price/performance/openness. In other words, we can be sure there is a bubble to the extent that open-weight models can successfully demonstrate that there is no moat. >Why doesn't China also worry about their own AI bubble inside the country? Because they haven't bet the farm on AI like the USA has.
- aurareturn 4mo agoTo weaken the stature of the USA on the global stage relative to themselves. Perhaps decrease US investment in AI and slow creation of some general AI superweapon I suppose. I think you are overthinking it. The reason why Chinese AI labs have to go open source is because they do not have the clout to freely expand to international markets. Therefore, in order to succeed and get attention, they are providing their models for free if you have the inference hardware. Well that's the position that these chinese firms are trying to convince us of, and they can convince us by undercutting proprietary models in price/performance/openness. I don't think these firms have said there is a US AI bubble and they're trying to pop it. Because they haven't bet the farm on AI like the USA has. They are. The only problem is that they can't buy Nvidia chips or EUV machines so they're bottlenecked.
- jamiedborin1 4mo agoI am the original author of the post - thanks for reading it! I think the future of open weights models will be similar to fabless chip design companies. There will be companies that can train models and they will licence those models to inference companies that manage the APIs. The inference companies need much less capital and the training companies dont need to divert resources from training to inference. Some of the Chinese model training companies are already doing this and licencing their models to inference providers.
- jingpostmedia 4mo ago[flagged]
- aurareturn 4mo agoThis is likely the case. I think people expecting companies to provide near-SOTA models for free forever are wrong. I think at some point, Deepseek or Z or other AI training companies will sell their models for fees. I can imagine buying an LLM model for $499 one-time payment for personal use. Maybe buying software and owning it will come back. Some will make you subscribe so you get the latest models as they release them. Of course, they will also license their models to inference providers like you said.
- alienbaby 4mo agoOn mobile, or at least on mine (pixel 10) using chrome, the graphs are unreadable and unusable, which is a shame as I'm quite interested in them and I don't have access to a pc at the moment. Would you be able to change that? They steal the scroll/drag touch and turn into a nightmare if zooming / unzooming, and are squashed and unreadable when they first render.
- try-working 4mo agoThis is exactly right.
- krater23 4mo agoHave no fear,after the bubble bursted, there will be more than enough cheap hardware for espacially this.
- profsummergig 3mo agoGodspeed.
- willmadden 4mo ago[flagged]
- mirekrusin 4mo agoClosed weight models are a result of philanthropy by some private investors.