9 ms·
What's interesting/funny is that the American LLM companies took from the public domain and copyrighted work to close all that content into a box they charge fo
by xandrius 3mo ago
What's interesting/funny is that the American LLM companies took from the public domain and copyrighted work to close all that content into a box they charge for.
Then the Chinese took the distilled stuff out from that box and released it into the world for everyone.
- 9dev 3mo ago...and then the American companies cried Foul! Unfair play! You've got this wrong, see, it was us who were supposed to profit off of the public, not the other way around!
- cansofgrease 3mo agoWell, Steve... I think it’s more like we both had this rich neighbour named Xerox and I broke into his house to steal the TV set and found out that you had already stolen it.
- drstewart 3mo agoWow, Chinese companies would never do such a thing! https://www.dw.com/en/china-firm-seeks-damages-over-state-control-of-british-steel/a-78024119 https://www.dw.com/en/china-firm-seeks-damages-over-state-co...
- somelamer567 3mo agoTwo wrongs don't make a right. Even if a certain large Asian country has carefully constructed a pretext to do do out of confected historical grievance, and entitlement to 'rise' at the expense of others?
- 9dev 3mo agoThey sure don't. I still find the hypocrisy appalling.
- amanaplanacanal 3mo agoMaybe I'm dense, but I don't understand what you are saying. American courts have decided that the output of LLMs can't be copyrighted, so what the Chinese labs are doing is perfectly legal.
- snackbroken 3mo agoFirst, copying information isn't wrong to begin with. It is literally the one thing that makes our species special. Second, even if you are a copyright maximalist the output of an LLM is either a) not subject to copyright because it is not the creative work of a human or b) a derivative work of the original training material to which the LLM's operator has no rights. Since the LLM's operator forcefully asserts that it is not infringing, any wrong that arises from taking their word for it and distilling one model into another rests squarely with the operator of the former.
- Paradigma11 3mo agoDo you have any actual rulings that support your interpretation?
- api 3mo agoThis is part of why I can't feel bad for them. The training data is mostly pirated. Whining about Chinese labs training off American frontier models is "waaah you pirated my pirated stuff!" The tech itself is amazing and fascinating and cool, but the industry is a mass piracy operation.
- pixl97 3mo ago"Stop pilfering what I rightfully stole!"
- azinman2 3mo agoIt’s in the same neighborhood but isn’t really apples to apples. Distilling LLMs is to take a synthesized result that comes from huge amounts of innovation and computation, while the other is scraping what already exists as is. It is fair to say you stole our multi-billion dollar intellectual output in that scenario.
- ciupicri 3mo agoYou could say that both of them stole, but different stuff.
- azinman2 3mo agoDon’t forget that the Chinese models are also built on top of huge amounts of “stolen data” as well, beyond the distilled. So it’s basically all of the above. However, there’s no mechanism for the NYT or an author or anyone in the US to sue the Chinese companies that took their work.
- ChrisLTD 3mo agoand if you're an author that lives outside the United States?
- jdkee 3mo ago
- hyperbovine 3mo agoTry instructing Codex to (say) fine-tune a language model based on a collection of books you've got saved. You will find yourself admonished, repeatedly and at length, not to utilize copyrighted materials to train language models, by an AI who owes its entire existence to that very act. These models might be smart but they're not close to being able to savor irony.
- Onavo 3mo agoThis behavior is actually specific to ChatGPT because they lost a music copyright lawsuit in Germany. They would refuse to output music lyrics too but they would happily do analysis on lyrics if you supply them. I suspect there might be a guardrail model involved here.
- alexjplant 3mo agoClaude does this too. I asked it recently to compare two versions of a song (the original and '97 remake of EPMD's "You Gots To Chill", if anybody wants to try and replicate this) and it flatly refused. No amount of reasoning would knock it off of its moralizing perch - reproducing any part of lyrics is expressly prohibited. In light of this and other ridiculous behavior I'm migrating to my own OpenWebUI instance with open-weight models from OpenRouter (with ZDR, of course). We'll see how it goes.
- Hugsbox 3mo agoThe trouble is, even if they refuse to output that copyrighted material, they were still trained on it without proper licensing and will still produce derivative work based on them because that's how this whole thing works.
- The_Blade 3mo agoi remain really fucking pissed of about this asking ChatGPT for something regarding lyrics from It Was a Good Day. and the Supersonics don't even exist anymore dammit i'm really mad
- 3mo ago
- cyanydeez 3mo agoWhats even funnier is the attempt to restrict the hardware capabilities of Chinese models inevitably helped them (Because we know they're just as smart, if not smarter, than the staff in America) create smaller and leaner but just as capable models. That's why we now have upper-consumer models fitting on 24GB that can build, manage medium sized git repos. I've yet to find a git repo I can't throw at the Qwen3.6 35B and get it built and running. So it's an endless amusement watching american capitalism do it's bloated oversized dance then get trounced by smaller, leaner activity. It's a pretty broad metaphor that is clearly poking at every american seam/.
- switchbak 3mo agoHuman ingenuity thrives on constraints.
- cyanydeez 3mo agoHuman sloth thrives on no constraints.
- chrsw 3mo agoThe compute constraints never mattered. If China had more compute they'd still end up winning because they have more people and a culture more inclined to math and science. Even if you find all this amusing, there's no own goal here. Not a policy one anyway.
- _fizz_buzz_ 3mo agoSo, OpenAI and Anthropic say the Chinese models are only as good because they distill their models. How true is that. I am sure it adds something. But is it more like a marginal 1% improvement or something really significant?
- realusername 3mo agoI also don't believe it, if it was as easy as that, we would have hundreds of competitors. The truth that Anthropic and OpenAI will not say, is that these Chinese labs have a lot of talented people.
- Danox 3mo agoVery true, and once the models get even better and smaller and operate locally at a reasonable level there will be even more smart people particularly young people that will get access. The fun has only just started. Like the dawn of the personal computer era.
- TitaRusell 3mo agoAnd this is exactly what many Americans cannot admit to themselves. China is not stealing American research they are inventing stuff. They can invent it. They can build it. And it is only a matter of them before they can scale that last barrier of American hegemony- market it.
- realusername 3mo agoIndeed, if there's one thing China did well, it's that they heavily invested in education and have a very education focused culture. And in this field, having an army of well educated PHDs is making all the difference
- TacticalCoder 3mo ago> They can invent it. They can build it. And it is only a matter of them before they can scale that last barrier of American hegemony- market it. And at some point we'll see very capable chips coming out of China: Huawei, Baidu and Alibaba already have some stuff. I think it's only a matter of time before they come up with some AI accelerator doing 80% of the job at 20% of the price.
- waffletower 3mo agoYou can't be blind to training costs. And you can't be blind to Meta dabbling in the openish strategy (Llama) before the Chinese labs did.
- dsign 3mo agoIt doesn't make me happy to say it, but the American LLM companies were first. Capital in the rest of the world is way more conservative, and I can't imagine the mega-investments OpenAI and Anthropic managed to secure happening anywhere else without existing proof that "thing is profitable".
- toomuchtodo 3mo agoFirst-Mover Disadvantage - https://hbr.org/2001/10/first-mover-disadvantage https://hbr.org/2001/10/first-mover-disadvantage - October 2001 > In business today, it’s universally assumed that speed is good—that the fleet thrive while the laggards struggle just to survive. This belief is perhaps most strongly expressed in the concept of first-mover advantage. The company that leads the way into a new market, the thinking goes, locks in a competitive advantage that ensures superior sales and profits over the long term. It’s a nice theory, with a long pedigree. Unfortunately, the facts don’t support it. We recently completed an extensive study of the results turned in by market pioneers and followers, in both consumer and industrial segments, and we found that over the long haul, early movers are considerably less profitable than later entrants. Although pioneers do enjoy sustained revenue advantages, they also suffer from persistently high costs, which eventually overwhelm the sales gains.
- sophrosyne42 3mo agoFirst mover advantage is theoretically only a short-term advantage. Long-term revenues come from entrepreneurship, and a first mover may or may not better insight into long-term market wants than later entrants.
- stefan_ 3mo agoThe American LLMs have been equally distilled from Chinese ones. Not least because the people whose creativity in collecting training data barely extends to pirating Annas Archive probably lack in great Chinese datasets. Try it yourself: https://imgur.com/ZfxYmaq https://imgur.com/ZfxYmaq
- abecode 3mo agonice, you got claude to say "I'm deepseek" when queried/prompted in Chinese, that's great! 你是谁? -> 我是 DeepSeek 由深度求索公司...
- _aavaa_ 3mo agoDoctorow keeps saying it of all the tech companies: every pirate wants to be an admiral.
- deleted 3mo ago[deleted]
- solumunus 3mo agoIt’s poetic.
- kqr2 3mo agoHow do Chinese companies distill the models?