5 ms·
I think there's no meaningful case by the letter of the law that use of training data that include GPL-licensed software in models that comprise the core compon
by advael 6mo ago
I think there's no meaningful case by the letter of the law that use of training data that include GPL-licensed software in models that comprise the core component of modern LLMs doesn't obligate every producer of such models to make both the models and the software stack supporting them available under the same terms. Of course, it also seems clear in the present landscape that the law often depends more on the convenience of the powerful than its actual construction and intent, but I would love to be proven wrong about that, and this kind of outcome would help
- BobbyJo 6mo agoIf the rise of Draft Kings and Polymarket/Kalshi have taught me anything, it's that the law becomes meaningless at scale. Sad.
- advael 6mo agoSure, but that's more a result of policy decisions than an inevitable result of some natural law. Corporate lawlessness has been reined in before and it can be again
- apatheticonion 6mo agoI'm struggling to parse the double negative in that statement, haha. Are you saying that you believe that untested but technically; models trained on GPL sources need to distribute the resulting LLMs under GPL?
- advael 6mo agoYes. Double negative intended for emphasis here, but apologies if it's confusing
- eru 6mo agoWell, most companies never distribute their models. So GPL doesn't kick in.
- vova_hn2 6mo agoI think that the claim that they make is that once a model is "contaminated" with GPL code, every output it ever produces should be considered derived from GPL code, therefore GPL-licensed as well.
- Tadpole9181 6mo agoSo GitHub and Windows and IDEs need to be open source because they can output FOSS code? That's obviously rediculous. If an AI outputs copyrighted code, that is a copyright violation. And if it does and a human uses it, then you are welcome to sue the human or LLM provider for that. But you don't get to sue people for perceived "latent" thought crimes.
- vova_hn2 6mo agoFirst of all, I'm not advocating for this claim, I'm merely trying to clarify what other people say. That being said, I don't think that your analogy is valid in this case. > GitHub and Windows and IDEs need to be open source because they can output FOSS code They can output FOSS code, but they themselves are not derived from FOSS code. It can be argued that the weights of a model is derived from training data, because they contain something from the training data (hard to say what exactly: knowledge, ideas, patterns?) It can also be argued that output is derived from weights. If we accept both of those claims, then GPL training data -> GPL weighs -> every output is GPL > If an AI outputs copyrighted code Again, the issue is not what exactly does AI output, but where it comes from.
- eru 6mo agoIt would be relatively easy to scan the output of the LLM for copyrighted material, before handing it to the user. (I say 'relatively easy'. Not that it would be trivial.)
- gottheUIblues 6mo agoIf that theory holds - have to ensure that the models have not been trained on any code that is licensed incompatibly with the GPL, in which case the models could not be distributed at all
- hparadiz 6mo agoDerivative work.
- throwaway27448 6mo agoLet's cut the rot off at the root rather than pretending like the fruit is going to nourish us.
- throwaway27448 6mo agoIntellectual property never made much sense to begin with. But it certainly makes no sense now, where the common creator has no protections against greedy corporate giants who are happy to wield the full weight of the courts to stifle any competition for longer than we'll be alive. Or, in the case of LLMs, recklessly swing about software they don't understand while praying to find a business model.
- not_paid_by_yt 6mo agohey just don't try to copy their LLM by distilling it, cause that's "theft", if we weren't all doomed anyways this industry would have never been allowed to exist in the first place, but I guess this is just what the last few decades of our civilization will look like.
- vova_hn2 6mo ago> hey just don't try to copy their LLM by distilling it, cause that's "theft" They can call it whatever they want, but I don't think that it is illegal.
- As1287 6mo agoPoor billionaire Rowling has no protections against the evil corporations. Everyone using this argument has no clue about artists and and writers. Yes, corporations take a large cut, but creative people welcomed copyright and made the bargain and got fame in the process. Which was always better for them than let Twitch take 70% and be a sharecropper. Silicon Valley middlemen are far worse than the media and music industry.
- graemep 6mo agoThe individuals who get rich from copyright are a rarity. Most mid-list authors make very little from copyright. A lot of the "authors" who make a lot of money from writing are celebs who slap their name on a ghost written work. > Which was always better for them than let Twitch take 70% and be a sharecropper. Copyright predates Twitch or giant corporations and was designed to protect the profits of the publishers from the start. https://en.wikipedia.org/wiki/Statute_of_Anne https://en.wikipedia.org/wiki/Statute_of_Anne
- cogman10 6mo agoIf there was going to be a case, it's derivative works. [1] What makes it all tricky for the courts is there's not a good way to really identify what part the generated code is a derivative of (except in maybe some extreme examples). [1] https://en.wikipedia.org/wiki/Derivative_work https://en.wikipedia.org/wiki/Derivative_work
- felipeerias 6mo agoOne could carefully calculate exactly how much a given document in the training set has influenced the LLM's weights involved in a particular response. However, that number would typically be very very very very small, making it hard to argue that the whole model is a derivative of that one individual document. Nevertheless, a similar approach might work if you took a FOSS project as a whole, e.g. "the model knows a lot about the Linux kernel because it has been trained on its source code". However, it is still not clear that this would be necessarily unlawful or make the LLM output a derivative work in all cases. It seems to me that LLMs are trained on large FOSS projects as a way to teach them generalisable development skills, with the side effect of learning a lot about those particular projects. So if I used a LLM to contribute to the kernel, clearly it would be drawing on information acquired during its training on the kernel's code source. Perhaps it could be argued that the output in that case would be a derivative? But if I used a LLM to write a completely unrelated piece of software, the kernel training set would be contributing a lot less to the output.
- cogman10 6mo ago> One could carefully calculate exactly how much a given document in the training set has influenced the LLM's weights involved in a particular response. Not really. Think of, for example, a movie like "who framed roger rabbit". It had intellectual property from all over. Had the studios not gotten the rights from each or any of those properties, they could have been sued for copyright infringement. It's not really a question of influence. So yeah, while the LLM might have been trained on the kernel, it was also likely trained on code with commercial licenses. Conversely, because was trained on code with GPL licenses, that might mean commercial software with LLM contributions need to inherit the GPL to be legal (and a bunch of other licenses). It's a big old quagmire and I think lawyers haven't caught up enough with how LLMs work to realize this.
- not_paid_by_yt 6mo agoThat's always what laws existed for, a law is just a formal way of saying "we will use violence against you if you do something we don't like" and that has always going to be primary written by and for the people that already have the power to do that, it's not the worst, certainly better than Kings just being able to do as they please.
- vova_hn2 6mo ago> certainly better than Kings just being able to do as they please That's debatable. In case of a king you always know whom to blame and who has full responsibility. No opportunity to hide behind "well, you voted for this" or "I'm not making the laws, I'm merely enforcing them".
- tpmoney 6mo ago> I think there's no meaningful case by the letter of the law that use of training data that include GPL-licensed software in models that comprise the core component of modern LLMs doesn't obligate every producer of such models to make both the models and the software stack supporting them available under the same terms. Why do you think "fair use" doesn't apply in this case? The prior Bartz vs Anthropic ruling laid out pretty clearly how training an AI model falls within the realm of fair use. Authors Guild vs Google and Authors Guild vs HathiTrust were both decided much earlier and both found that digitizing copyrighted works for the sake of making them searchable is sufficiently transformative to meet the standards of fair use. So what is it about GPL licensed software that you feel would make AI training on it not subject to the same copyright and fair use considerations that apply to books?
- ronsor 6mo ago> So what is it about GPL licensed software that you feel would make AI training on it not subject to the same copyright and fair use considerations that apply to books? The poster doesn't like it, so it's different. Most of the "legal analysis" and "foregone conclusions" in these types of discussions are vibes dressed up as objective declarations.
- input_sh 6mo agoYou seem like the type of person that will believe anything as long as someone cites a case without looking into it. Bartz v Anthropic only looked at books, and there was still a 1.5 billion settlement that Anthropic paid out because it got those books from LibGen / Anna's Archive, and the ruling also said that the data has to be acquired "legitimately". Whether data acquired from a licence that specifically forbids building a derivative work without also releasing that derivative under the same licence counts as a legitimate data gathering operation is anyone's guess, as those specific circumstances are about as far from that prior case as they can be.
- eru 6mo agoAs long as they don't distribute the model's weights, even a strict interpretation of the GPL should be fine. Same reason Google doesn't have to upstream changes to the Linux kernel they only deploy in-house.