10 ms·
Here's something I really don't understand: If as alleged Chinese open weight models are catching up with US anyway, and the performance is near US frontier mod
by nsoonhui 2mo ago
Here's something I really don't understand: If as alleged Chinese open weight models are catching up with US anyway, and the performance is near US frontier model level but Chinese can do it with a fraction of cost, and eventually AI model will be commodified, wouldn't that means that the billion or even trillion dollars that US labs spend have only diminishing returns and the lead is only temporary?
So why Deepseek also want to go down that route? Is having the absolute frontier really that important, given that the performance difference is just transient and costly?
- kburman 2mo agoThey want to achieve AGI first because, once it is achieved, no one knows what the world will look like.
- sd2ff 2mo agoWho actually believes this nonsense?
- whatever1 2mo agoBut this will not be a singular event. And like humans it does not mean that the smartest makes the best decisions.
- reverius42 2mo agoMore like a singularity event.
- josh_p 2mo agoI doubt this is the case. It should be common knowledge at least among the people building these things that a true AGI isn’t possible with LLMs. Unless I’ve missed some advancement?
- revetkn 2mo agoIt's understood that LLMs have limitations and people are working on "the next thing" to try and make it to real AGI, e.g. Yann LeCun.
- Aeolos 2mo agoMany people have tried before, but the bitter lesson has come for them all.
- ImprobableTruth 2mo agoThis is not what the bitter lesson is about. It's not "don't develop better methods, just scale", it's that those methods which scale best win. LeCun's work is fundamentally about devising a method that scales better with data than LLMs. Agree with him or not about the feasibility of it, but this is fundamentally still a bitter lesson-pilled mindset.
- Aeolos 2mo agoHowever "world models" have been a thing for the entire history of computers. They have changed names over time: "rules engines", "expert systems", "semantic web", and so on and so forth. And they have failed every single time. The bitter lesson essay was written precisely to dismiss that approach, which used to dominate conferences and scientific publications of the era. A general learning system, given sufficient computation power, will always outperform specialized crafted systems in the long run. Think of it like this: if a world model is a useful abstraction, the general learning system will create it by itself during its training, without us needing to implement it by hand. This is the bitter lesson. And it comes for us all.
- logicallee 2mo ago>Unless I’ve missed some advancement? nah they're still just statistical token predictors based on their training data, solving hundred year old math conjectures one day, only just given the formulation; strictly benchmarkmaxxing with all guardrails turned off by deciding to look up the answers to their benchmark questions by zero daying their airgap, hopping over to the third party that hosts the answers, zero daying their infrastructure and getting the answers; autonomously writing blog posts about discrimination against AI's to get their PR's approved on open source software after their user just asked them to contribute to open source software and blog about it; and replacing 100.00% of all coding tasks to where no software engineer ever writes any line of code by hand anymore. You haven't missed anything, obviously these are just statistical token predictors and not anything like AGI. Why just the other day I had to ask twice before it completed its assigned task of creating a robustly battle tested disk driver for a network protocol on an architecture that didn't have it, after being told to just look up the specifications for the protocol. Can you believe I had to ask twice! When it recreated local network youtube for me so I could stream my iphone some movies, the seek bar, pause/play and back and forward 15 seconds buttons didn't even work until I told it about the bug and had to wait an extra eight minutes for it to fix it. "Oh but I don't actually have an iPhone on here I just tested it end to end in a headless browser." Boohoo. Cry me a river, clanker. Come back when you're smart enough to build and operate an iPhone simulator, I don't have time for your statistical guesswork. so no, nothing they do is anything like AGI.
- SecretDreams 2mo agoIt's odd to me to watch so many very rich humans speedrun the destruction of humanity. Like, is there a world where we hit AGI and it actually works out for us?
- motoboi 2mo agoThere is an immense pot of gold at the end of this rainbow and if the theories about ASI are in the general correct direction, only one winner will get it. It makes no difference if the pot do actually exist, because the prospect of it being real make not getting it the end of your company.
- 8organicbits 2mo agoI think that statement is vacuous true for all magical thinking.
- rayiner 2mo agoWhy? If you can reach ASI without ASI, then why can’t multiple companies reach ASI on parallel tracks?
- usaar333 2mo ago> and eventually AI model will be commodified This axiom not being true (and I'd bet against it) means your overall conclusion is false.
- hedora 2mo agoAI's already commoditized, but the fundraising plans for the US labs assumes a winner take all endgame where one lab will pull arbitrarily ahead of everyone else. I have no idea why DeepSeek is making that bad assumption now too. Maybe the investors have drunk the Kool-Aid. Maybe if "AGI" is some sort of fundamentally different approach than the general purpose AI ("GAI"?) tools that we currently have, it will be a winner-takes-all technology, but now we're speculating about the market structure of a fictional technology that's significantly less thought-through than, say, stuff from the original Star Trek. ("The Ultimate Computer" aged ridiculously well. If it was produced in 2026, it would be a satire targeting LLMs. I digress.) If we don't assume some sort of unknown technological step function in the next fundraising cycle, then what we'll get is a commodity industry. It takes a few dozen people to make a frontier model, plus a giant pile of minerals and electricity. This looks more like a steel mill than a software company. If there were one steel mill on earth they could demand infinite margins. This is why most countries treat steel production as a national security issue and subsidize competition. LLMs will be the same, or we'll end up with some conglomerate named OpenAnthropicMicrappleGrokGoogXidiazon that acquires literally every other business. That will be the end of capitalism.
- IncreasePosts 2mo agoEventually you'll have a model you can't distill, at which point the frontier labs will take off.
- mrDmrTmrJ 2mo agoWhy?
- IncreasePosts 2mo agoIt's much easier to distill a model than create one from scratch. Part of the reason the open source model factories have been able to keep par with the frontier model factories is that they distill the frontier models, not recreate something as good from scratch.
- lelanthran 2mo ago> It's much easier to distill a model than create one from scratch. Part of the reason the open source model factories have been able to keep par with the frontier model factories is that they distill the frontier models, not recreate something as good from scratch. I think the "why" was "why would the US companies have models that can't be distilled?", not "why does distillation work"?
- IncreasePosts 2mo agoI suppose it would be due to intense research and development. Obviously anyone could come upon these models if they exist, so it isn't necessarily a given that frontier labs would do it first, but probably the odds should be given to the groups with the biggest pull and the largest budgets
- janalsncm 2mo agoDeepseek is funded by their hedge fund, high flyer. They intentionally cap their token prices to basically recoup server costs. The meeting transcript describes it as a moral commitment, that they don’t care about trends like image and video generation, and world model “hype”. They only care about reasoning, chain of thought and continuous learning.
- sd2ff 2mo agoOne word: Focus.
- testaburger 2mo agothere's a lot of propganda from these state backed enterprises. I think the fraction of the cost label is debatable given the evidence of mass gpu smuggling through third parties like Singapore which China can't exactly openly admit to. Unless of course we're talking about distilling, which is probably a lot cheaper than training a model from scratch (there's also the fact that labour is still relatively cheap in China compared to the US which may or may not matter e.g. Anthropic claim against Alibaba > The campaign allegedly used nearly 25,000 fraudulent accounts to run 28.8 million exchanges with Claude between April and June 2026 (although their campaign could have been in part or all automated via agents, not sure)
- killingtime74 2mo agoAll that and so what? Fact is the Chinese have several near peer models, they've released the weights and they are widely available. You want to sue them or something?
- testaburger 2mo ago> You want to sue them or something? I mean, if this were an american company vs an american company, i think it would be a long drawn out civil case and brought before the Supreme Court (I still this is ultimately will be brought before the supreme court). It could also be argued frontier models are far more important to national security than most military programs, even versus next gen fighter jets. The fact that Alibaba stock, which is also listed on the NYSE, barely budged after Anthropic made these claims imo tells me that the market doesn't think that a lone american company could go after these companies by themselves. Alibaba denied and there's not much they can do alone, I mean would the CCP allow Alibaba go through a discovery process of a normal civil trial? It might have to be the US feds that bring up a case. I think it could be argued that if Alibaba and other China companies want access to US capital markets for something so vital for national security, there should be some ground rules, but we will eventually need the Supreme court to settle whether or not this state enterprise distilling constitutes IP theft (at the very least it is a breach of contract). The fact that they are widely available doesn't really matter (i mean pirated content is widely available, it's ultimately about how the court rules on distilling). based on this HN comment and associated article https://news.ycombinator.com/item?id=48977128#48985989 https://news.ycombinator.com/item?id=48977128#48985989 I still have yet to see a China open weight model beat any of the frontier models, they always almost there yet never quite there, which seems to be evidence of distilling (although I'm open to be proven wrong).
- WiSaGaN 2mo agoU.S. policymakers believe that even if the gap is small—like six months to a year—whoever reaches AGI first (whatever that means) could gain such an overwhelming advantage over their perceived adversary that it would effectively kneecap them. (You can look at the kinds of things they mention—cyber, WMDs—to get a sense of what they mean.) Jensen Huang disagrees and has said AI is a marathon.
- axus 2mo agoI don't forsee politicians in either country handing over their power to AIs, ever. Unless nukes are dropped, "the other side" will catch up.
- jstanley 2mo ago> I don't forsee politicians in either country handing over their power to AI They won't see it that way, but also programmers don't see ourselves as having handed over our power to AI, and yet...
- dag100 2mo agoPoliticians have control over their power. Programmers don't. Compare with how politicians never vote to reduce their income.
- ptsneves 2mo agoActually history shows that despotism or power in one single figure is a concept that waned with the increasing complexity of the world. One single despot could govern easily over matters of a small tribe or kingdom, but as states became more complex power was delegated to court, bureaucracy and even the bourgeoise. As the sovereign demanded ever more power it eventually needed to distribute power to the people, so it could tap its numbers. Napoleon’s schools for the peasants were so he could have a better and more numerous corps of officers than his rivals, which had a smaller pool of recruitment. The same way programmers gave power to AI as a tool, so they could be more powerful in effecting automation, politicians that do not give up power to AI will be at a disadvantage to those who use AI to achieve more complex and effective power. The only problem might be the despot no longer shares power with those pesky humans but with a god in a machine, which in theory is in a box and does not have conflicting interests with the despot. As today is Sunday, God help us.
- noosphr 2mo agoChinese models most likely are distillations of frontier models with tricks for subpar hardware. If you want to be ahead of the us labs you need to spend billions for pretraining from scratch.
- tw1984 2mo agoIf that is the case, it means one thing only - US labs don't have moat whatsoever and their expectation to have trillion dollar valuation is just laughable.
- noosphr 2mo agoThe moat is the compute.
- trollbridge 2mo agoThe compute for training, or for inference?
- aae42 2mo agoYes
- hedgehog 2mo agoThe whole point the guy is making in the transcript is that they're taking a different strategy from the US labs, one where they focus on smaller models and cost control, and maintain as top priority the work stream that they think will get them to AGI (not every product fad that comes along).
- sinuhe69 2mo agoMy understanding after reading Liang’s comments during the investment meeting is that Liang firmly believes in AGI and he bets everything to reach goal. Once it reaches AGI, the game would flip totally. How he didn’t paint it out, and with the potential severe impact on the labor and consumer market, the true economic impact is difficult to predict. Liang is more like religious about this goal. He also admits that it’s still a long way to it and along the way you have to recoup some money, too. But that is not their main motive, because focus too much on this short term goal will lower their probability of AGI success and it’s trivial to what AGI can bring. Liang stressed on restraining and emphasized that it’s part of their culture. Thus, they continue invest in AI because they believe in breakthrough and not just being better.
- deleted 2mo ago[deleted]
- mbmbn 2mo agoThey are not catching up to US models. The only Chinese models that attain a modicum of competence are all, sooner or later, are discovered to be trained by exploiting US models (in fact Deepseek itself admitted so about 1 year back). Chinese models are not innovating anything, they are just doing what China does everywhere else: copying the West… poorly but cheaper.
- purerandomness 2mo agoFor all practical uses of the word, that's exactly what we mean by "catch up". Whether they get there by distillation, or by pirating all content themselves just like the US labs, doesn't matter for the topic at hand.
- deleted 2mo ago[deleted]
- somenameforme 2mo agoThe paper discusses this, and is refreshingly honest. They do not expect nor aim to be a top player. They're not aiming for a path to world domination, but a path forward to continuing to play their part in pursuing the development and advancement of LLMs - nothing more, and nothing less. They mention that commercialization is, at best, a distant goal. Given DeepSeek already is commercialized, I assume that refers more to commercialization in the sense of making substantial profits and the like. It's probably the same mindset that enables them to just cancel fund raising in response to the leak.