4 ms·
So what is the issue here? Distilling is still fair, on the same level like Anthropic scraped copyright protected material for their training. So here robbers
by madduci 2mo ago
So what is the issue here? Distilling is still fair, on the same level like Anthropic scraped copyright protected material for their training.
So here robbers are blaming robbers?
These claims are just pointless, everytime
- xnoto 2mo ago++
- make3 2mo agoIt's about the claim of whether these companies could develop a similarly powerful model without larger companies building their own first, which is an important point, and it's likely not the case. It's also about the larger companies explaining why they can't be as efficient, of course they can't, they're not just ripping the outputs of another model that someone else invested billions to train.
- PaulHoule 2mo agoSimply knowing it is possible to do something makes it easier to do.
- cindyllm 2mo ago[dead]
- doctoboggan 2mo agoYeah agreed, from one standpoint I couldn't care less that they did a "distillation attack", but I am interested in knowing if China is able to develop open weight frontier models without the prior existence of a huge model to distill from.
- IncreasePosts 2mo agoWhy would that matter? OpenAI or whatever frontier lab couldn't have built their frontier models without the entirety of humanity unknowingly developing their training set for 5000 years. It would be one thing if Moonshot was breaking into OpenAI servers and stealing trade secrets, but the only thing they are doing is looking at the output of the program, which is exactly the service that OpenAI offers. So, at best, this is a ToS violation. Sucks for the frontier labs I suppose, but live by the sword - die by the sword.
- mrhottakes 2mo ago> they're not just ripping the outputs of another model that someone else invested billions to train. True, they're simply ripping the inputs that humanity invested thousands of years and trillions of dollars to produce.
- make3 2mo agono argument from me here
- xcf_seetan 2mo ago> they're not just ripping the outputs of another model that someone else invested billions to train. If they payed for inference, doesn't they own the output? So if I pay for a model to generate code, isn't that code mine to do with it whatever I want? Just curious.
- make3 2mo agoNot arguing for the morality of it, but if we're going by the law because that's what you're using in your comment ("don't I own" which only matters wrt the law), then you explicitly accepted a Terms of Use which excludes distillation as a use case. Now of course they themselves trained on the whole Internet for free, etc.
- dgellow 2mo ago> on the same level like Anthropic scraped copyright protected material for their training. I see no problem with distillation, on the other hand the complete dismissal of copyright by AI labs is pretty bad, I don’t think we should put them at the same level
- mrtesthah 2mo agoThe amount of original, copyrightable and trademarkable IP actually created by the AI labs themselves is dwarfed by their staggeringly vast infringement activities.
- SR2Z 2mo ago> on the other hand the complete dismissal of copyright by AI labs Courts keep ruling over and over that an LLM trained on copyrighted works qualifies as a transformative work and is therefore fair use. They don't have to dismiss copyright law, this has always been allowed. The only thing they get in trouble for is pirating the works to get their hands on them.
- Diogenesian 2mo ago"Keep ruling over and over" is way too strong. There have maybe been two rulings, nothing nationally binding, and most of the litigation is still ongoing. In particular, last I checked OpenAI and Microsoft are still badly threatened by the NYT lawsuit: https://law.justia.com/cases/federal/district-courts/new-york/nysdce/1:2023cv11195/612697/514/ https://law.justia.com/cases/federal/district-courts/new-yor... https://www.cnet.com/tech/services-and-software/publishers-openai-copyright-lawsuits-nyt-ziff-davis/ https://www.cnet.com/tech/services-and-software/publishers-o... This will have to wait for the Supreme Court. OpenAI and Microsoft 100% deserve to lose, even without OpenAI allegedly hiding evidence.
- asadotzler 2mo agoNot to mention the cases where the AI labs would have lost in court so bailed and settled for billions. Just this week, Anthropic agreed to pay $1.5B in a settlement to avoid losing a pretty cut and dry case.
- Matl 2mo ago> So what is the issue here? The issue seems to be the US only likes competition when it is winning. Markets in Asia are meant for cheap labor and resources, they're not meant to actually compete. /s
- catigula 2mo ago[flagged]
- ceejayoz 2mo agoEverything you just said describes the major American AI providers. Anthropic just settled a $1.5B suit over it!
- fwip 2mo ago[flagged]
- soperj 2mo ago> a bad actor that leverages Ip theft wholesale. It's like they've read the history of the US and how it got to where it is in the first place.
- nickphx 2mo agoOh, ok. How would you describe how the "frontier us companies" acquired the data used to form their models?
- rickydroll 2mo agoStealing IP is how American industry got started. Goose: gander, pot: kettle. It is what built and sustains the movie and music industries. See: work for hire and 100+year copyright length The tech industry: see: copyright and patent assignment from discoverer to corporation. I know that corporations forcing me to assign patents and copyright to them was an incentive to take published works from "software practice and experience" and other technical journals, use them as the core of my work, and disclose that source to the company I worked for. Didn't stop them from applying for patents, however. I think the discussion of copyright needs more refinement. We need to separate the discoverer's need for acknowledgment of development effort from the rent-seeking core of copyright.
- petilon 2mo ago[dead]
- Lalabadie 2mo ago"You are trying to kidnap what I have rightfully stolen!"
- sciencesama 2mo agothe whole AI is just internet distilled !!
- azinman2 2mo agoExcept it’s not just a dump of the internet, which Moonshot also did themselves (and probably used even more pirated content as laws in China are different without any recourse for the entire world). I don’t know why this is so unclear to folks.
- trollbridge 2mo agoChinese IP law is actually quite solid. You do have to register your trademarks and copyrights properly in China, and then lawsuits have to be filed appropriately according to Chinese law. Is that a problem?
- JKCalhoun 2mo agoI'm by no means taking the side of the AI companies, but it's possible that Anthropic "added value" to the data they harvested. Stealing that does seem kind of uncool. Regardless, it was always inevitable—will continue to happen.
- oliculipolicula 2mo agoValuation is hard to perform when it's deep inside a black box. Ther "API" may be easier to evaluate. The problem with this angle is that Moonshot is actually producing _better_ value from Anthropic's blackbox. Technically, providing better value from your competitor's private holdings could be theft (of trade secrets), but might it also be fair use? "Schrodinger's IP" be damned. I don't think the 1.5B settlement has resolved this. The 2 cases need to be merged!
- mrhottakes 2mo agoSo as long as Kimi added value to Fable, it's fine? Sounds good.
- knollimar 2mo agoMoonshot can prove they added value with their paper. Where's Anthropic's proof?
- 6gvONxR4sf7o 2mo agoI wonder how the "added value" argument can apply to Anthropic and not to the Kimi team. If Claude's value is that you don't have to pay a team of slow expensive subject-matter experts, and Kimi's value is that you don't have to pay Claude, it just seems like the same thing.
- orangecat 2mo agoDistilling is still fair I generally agree, in the same sense that it's "fair" for the US and China to spy on each other. It's not a moral outrage, but it is something that the targets can and should try to prevent.
- archagon 2mo agoOutrageous only to the died-in-wool corpocrats.
- softwaredoug 2mo agoIf they did this in the US they would almost certainly be sued. Meta, for examples, doesn’t want employees to use Claude Code due to distillation risk.
- trollbridge 2mo agoIt turns out U.S. law doesn’t have jurisdiction across the entire world, nor does Anthropic and OAI’s rather blatant attempt to buy government influence.
- softwaredoug 2mo agoWell its not law. Its more terms of service. For example, if OpenAI / Anthropic were actually open, other US labs could be building near-frontier open weights models by distilling off OpenAI / Anthropic. But because US companies don't want to be sued, US labs who obey terms of service, will be at a disadvantage to Chinese peers. Maybe US labs need to just not care and distill from OpenAI / Anthropic anyways?
- deleted 2mo ago[deleted]