3 ms·
I am saying this probably is "silly behavior by a government" and it is a milestone that points towards what the future may look like. Why can't it be both? It
by libraryofbabel 4mo ago
I am saying this probably is "silly behavior by a government" and it is a milestone that points towards what the future may look like. Why can't it be both?
It's easy to wave this aside as the current administration playing political games. But I don't think there is any reason to assume that the current era of open availability of models is going to continue indefinitely. Do you think that Chinese labs will continue to release open models forever, even why they get to the level that Mythos is at now, and beyond? And do you think that a competent US government would have no interest in regulating and restricting model access in 2 years time, assuming that model capabilities continue to improve? I think we bias towards thinking the status quo is the norm and will continue, but this news invites us to question that assumption and think about different ways the future could go.
- gpm 4mo ago> Do you think that Chinese labs will continue to release open models forever Yes. I think the Chinese government either already has, or will soon, grasp that if they train the models that people use they dictate what people believe (at least around the margins where that's malleable), and they will happily throw resources at that. And simultaneously that the only way they can actually get everyone to use their models is if it's possible for us to run them on our own hardware. (This isn't exactly a utopian view of the future)
- tw1984 4mo ago> I think the Chinese government either already has, or will soon, grasp that if they train the models that people use they dictate what people believe (at least around the margins where that's malleable), and they will happily throw resources at that. that doesn't require the model to be SOTA, it can be just a compact model capable of running on some inexpensive hardware. that is vastly different from SOTA models like Mythos which can potentially disrupt lots of things.
- strangegecko 4mo agoOf course it requires SOTA, people will always choose better models over some compact thing that is obviously more limited. You can't control the truth with models nobody wants to use.
- columnarx3 4mo agoPeople choose SOTA right now because of the heavily subsidised model subscriptions. People aren't going to pay 20x the price for a model that's maybe 10% better.
- ezst 4mo agoAnd the fact that "better" is highly subjective and domain/task/vibe-specific
- adrianN 4mo agoWhy do I want the model I use for coding to know Shakespeare or vice versa?
- Jare 4mo agoBecause you communicate with it using natural language and real-world references and descriptions of what you want, you use emotion and emphasis (especially when re-prompting), you use examples and illustrative stories and common expressions. Understanding and interpreting all of that and replying in kind, to some degree, requires a large body of non-computation, cultural knowledge, or else the prompts are just meaningless words, and the replies will look like compiler output.
- adrianN 4mo agoThat sounds intuitively true, but I’m not convinced that it is actually the case. I don’t think we know enough about neural network training to say what training and how many parameters are necessary for what kind of performance on which tasks. To me it looks like we currently guess that more is better and try to throw as much compute and data at the problem as is economically feasible. There is little incentive for companies to invest into small model research since their moat is huge models that require special hardware to run.
- jychang 4mo agoThis is going to age very poorly when the best Chinese labs ALREADY just started not open sourcing their models. Qwen 3.7 is not open source; previous Qwen versions would have open source releases, but Qwen 3.7 plus does not. The second best Chinese model, Minimax M3, is testing the waters by taking longer and longer between “model release” and open sourcing it. This time, they spent 2 weeks after release before open sourcing it. There’s also a lot of rumors of GLM and Deepseek not open sourcing future models. It’s pretty obvious that you cannot take Chinese models as open source for granted, they’ll be closed source soon.
- ls612 4mo agoThe main reason the Chinese labs are releasing models as open weights is because they don't have the compute necessary to provide all of the inference. For the US frontier models something like 80-90% of the lifetime compute required for the model is inference rather than training. China wants to shepherd as much of their limited compute as possible towards training to keep up in the race.
- londons_explore 4mo agoWith nearly everyone using inference accelerators, the pool of hardware is no longer shared between training and use.
- SubiculumCode 4mo agoNo, they are open sourcing them because they don't have another play, being second/3rd tier lans
- Slartie 4mo agoI think the main reason is to minimize the market for closed-source models from US companies. China knows that doing what Anthropic/OpenAI/Google/... are doing is impossible for them. No one outside of China in any sane condition will send their data to compute farms IN CHINA like people currently do with US-based frontier models. Even if they could muster the inference power. Hence they do the second-best thing possible to attack the dominance of the US-based corporations: reduce their moat by open-sourcing models that are not fully equal, but practically useful and good enough for easily 90% of typical tasks people use agents for in their daily lives. But way cheaper to run. As long as this arms race in AI continues, China as "number two" will have some incentive to continue open-sourcing models. But of course the US government might force a change if they continue to enforce limited public access to new frontier models - there is no market to minimize if a model is not allowed to be publicly available.
- nine_k 4mo agoThe US administration restricting the use of US-trained models is one of the best gifts it could make to the Chinese LLM producers, and to the PRC government.
- dozerly 4mo agoThis entire administration is a gift to everybody but the US. It’s either in service of Russia, China or whoever is willing to pay Trump the most.
- rjzzleep 4mo agoChinese have a nickname for Trump. 川建国. Trump the nation builder(meaning China). But Biden actually continued most of Trumps policies.
- Der_Einzige 4mo agoI won’t forgive Biden for not reversing more of trumps policies, especially immigration Between RBJ refusing to step down, Biden not reversing immigration policy, and Biden refusing to step down in the primary until too late, he’s going to go down as a poor president in the history books - even if he wasn’t a bad dude or even bad in terms of policy.
- FpUser 4mo agoHe was getting senile. What did you expect. There must be age limit for rulers
- Der_Einzige 4mo agoTrump was also getting senile before they attempted to assassinate him. Hatred of his enemies gave him another 5 years of energy. Very frustrating, because he absolutly was doing word salad nonsense like this regularly before someone tried to shoot him: "Look, having nuclear — my uncle was a great professor and scientist and engineer, Dr. John Trump at MIT; good genes, very good genes, OK, very smart, the Wharton School of Finance, very good, very smart — you know, if you’re a conservative Republican, if I were a liberal, if, like, OK, if I ran as a liberal Democrat, they would say I'm one of the smartest people anywhere in the world — it’s true! — but when you're a conservative Republican they try — oh, do they do a number — that’s why I always start off: Went to Wharton, was a good student, went there, went there, did this, built a fortune — you know I have to give my like credentials all the time, because we’re a little disadvantaged — but you look at the nuclear deal, the thing that really bothers me — it would have been so easy, and it’s not as important as these lives are — nuclear is so powerful; my uncle explained that to me many, many years ago, the power and that was 35 years ago; he would explain the power of what's going to happen and he was right, who would have thought? — but when you look at what's going on with the four prisoners — now it used to be three, now it’s four — but when it was three and even now, I would have said it's all in the messenger; fellas, and it is fellas because, you know, they don't, they haven’t figured that the women are smarter right now than the men, so, you know, it’s gonna take them about another 150 years — but the Persians are great negotiators, the Iranians are great negotiators, so, and they, they just killed, they just killed us, this is horrible." - Donald Trump, 2016
- close04 4mo agoYou don’t need the cutting edge to influence people’s opinion. “Export LLMs” to the rescue.
- vintermann 4mo agoThere's also the Meta motivation, that even if you don't get the control you would like from releasing a model, it may still be worth it to at least deny others that control. I'm sure that matters even more to China vs. the US than it mattered to Facebook vs. Google.
- baq 4mo ago> > Do you think that Chinese labs will continue to release open models forever > Yes. holy shit the naivete of HN nowadays.
- spiralpolitik 4mo agoThere is no moat in the model and by making the them open, it’s hard for one to be established when the free models are “good enough”. OpenAI and Anthropic are both hamstrung by this. Anthropic does have the better chance of surviving.
- deanishe 4mo ago> Why can't it be both? Is the government going to fund all further development? Hard to imagine investors continuing to throw billions at products they aren't allowed to sell.
- CraftingLinks 4mo agoWhy wouldn't they? They see this technology as a military asset now.
- VBprogrammer 4mo agoHonestly, with the caliber of people who currently comprise the US administration; leaving the whole thing to Openclaw and some new fancy model might not be the worst idea.
- deleted 4mo ago[deleted]
- layer8 4mo agoTrump and friends are only interested in investments they can personally make money from.
- locknitpicker 4mo ago> I am saying this probably is "silly behavior by a government" and it is a milestone that points towards what the future may look like. Why can't it be both? Here is why it's unlikely this is anything other than "silly behavior by a government": - some benchmarks show GPT-5.5, Gemini 3.1, and even Claude Opus outperforming Claude Fable, and yet it's Fable which is restricted. - some benchmarks still show the likes of Kimi 2.5 outperforming any Claude model, and DeepSeek is getting equivalent scores (a few tenths of a percent difference) > Do you think that Chinese labs will continue to release open models forever (...) That's immaterial to the discussion. Even if China forced Chinese labs to restrict access to all models, the truth of the matter is that Trump's administration to restrict access to US-based models does not prevent others from having access to models that are as capable or even better. So what's exactly the point of this?
- rileyphone 4mo agoAll that says is some benchmarks aren’t worth the tokens it takes to evaluate them. Mythos is clearly capable of finding zero days other models can’t, and Fable is close enough to be lumped with it.
- mullingitover 4mo ago> Mythos is clearly capable of finding zero days other models can’t I'm unconvinced that this is anything more than proof of work and marginal improvement that other models will catch up with, perhaps as early as to next week. Lots of other current-gen models will find vulns that can be chained together if you're willing to burn enough tokens on the task, and Fable is an absolute token incinerator.
- solumunus 4mo agoYou’re completely overrating these benchmarks and it’s landing you at a nonsense opinion. Just actually use the models and you will see that the gap is significant.
- irthomasthomas 4mo ago
- JohnBooty 4mo agoYeah, there’s been a lot of debate about this on r/localllama — will there be a steady supply of new free/open models in the future? And if not, can we simply keep augmenting “stale” models with new knowledge to keep them useful? I’m on the pessimistic side of things on both questions. As for the second question, obviously stale models can be augmented to an extent but it’s nowhere near a substitute for new knowledge being fully baked directly into its training.