7 ms·
Copyright violation is not stealing, and training is not copyright violation (it's already been ruled as fair use, multiple times).
by 1gn15 1y ago
Copyright violation is not stealing, and training is not copyright violation (it's already been ruled as fair use, multiple times).
- inglor_cz 1y agoI think the concerning problem is when the LLM reproduces some copyrighted code verbatim, and the user doesn't even stand a chance to know it.
- 1gn15 1y agoYes, but that's not what the grandparent comment was talking about.
- isodev 1y agoIf I’m the grandparent comment, it was a big part of what I mean. Stolen/Unknown content goes in for training, verbatim or very close “inspired by” code comes out and there is no way to verify the source - “violation as a service”.
- fluidcruft 1y agoVerbatim dumping is one thing but otherwise this seems closer to the issue of plagiarism than copyright. If someone studies the Linux kernel and then builds a new kernel that follows some of the design decisions and idioms that's not really copyright infringement. The bigger issue (spiritually anyway) seems to be the need to develop free software LLM tools the same way FSF needed to develop free compilers. That's what's going to keep users from being able to adapt and control their machines. The issue is more ecological that programmers equipped with LLM are likely much more productive at creating and modifying code. Some of the rest seems more like saying that anyone who studies GCC internals is forever tainted and must write copyleft code for life which seems laughable to me. Again this is more a topic of plagiarism than copyright which are fairly similar but actually different and not as clear cut.
- isodev 1y ago> more a topic of plagiarism than copyright You’re right, in the context of a technical legal interpretation they’re different. In the context of right or wrong, they amount to the same. > anyone who studies GCC internals LLMs are not a someone, they’re more like … the indigo printout of some text or design, you then use to make a scrapbook to be mass produced for profit. Very different situation. When the AI bubble pops, I hope we will have some equalisation back to something more ethical.
- tjr 1y agoLLMs are not a someone This is in line with my disagreement over the fair use rulings. Most people who published works that have been used to train AI systems, created those works and published them for other people to consume and benefit from, not for proprietary software systems to consume and benefit from. The existing licenses and laws did not account for this; nobody was anticipating it.
- fluidcruft 1y agoI don't know... there's a pretty clear difference between copyright and a say a utility patent or trade secret. The right and wrong in FSF isn't about labor, it's about control over machines and ability to modify. Free software has never tried to control the community using patents and trade secrets and in general are rather hostile to them. In fact its fairly contemptuous of copyright and uses it from a purely utilitarian perspective. And frankly FSF is not opposed to commercial software. They're opposed to users being unable to modify machines and software that they are using. That's the core of the ethics. See the origins in that damn printer firmware RMS did battle with. But I also disagree in general about LLMs. LLMs are statistical text models but the general concept of what if there were an "AI" that wasn't a LLM and was trained on open source software is the same at the end of the day. I think whether or not LLM are intelligent or equivalent to humans is a red herring. There's no reason to not consider the implications of machines that are indistinguishable or even superior to human programmers. Particularly if we're discussing ethics getting lost in implementation details seems like a distraction and then all the derived ethics gets thrown out after the next innovation.
- CamperBob2 1y agoWhen that happens, it's because the code was trivial enough to be compressed to a minuscule handful of bits... either because it literally is trivial, or because it's common enough to have become part of our shared lexicon. As a society, we don't benefit from copyright maximalism, despite how trendy it is around here all of a sudden. See also Oracle v. Google.
- thesz 1y agoQuake's sqrt approximation is not trivial and is not common. [1] https://www.reddit.com/r/programming/comments/oc9qj1/copilot_regurgitating_quake_code_including_sweary/ https://www.reddit.com/r/programming/comments/oc9qj1/copilot...
- CamperBob2 1y ago(Shrug) It's a math trick, documented by Abrash among others and very heavily discussed on forums such as this one. And it didn't originate in the Quake codebase. Like much IEEE754 hackery, it goes back to the father of IEEE754 himself, William Kahan. Nobody benefits from a law that says that LLMs can't regurgitate the Quake sqrt() approximation. If that's what the law actually says, which it isn't.
- isodev 1y agoNot really, only a handful of authorities have weighed on that and most of them in a country where model providers literally buy themselves policy and judges.
- matheusmoreira 1y agoYeah, copyright infringement isn't stealing, copyright shouldn't even exist to begin with. I just think it's especially asinine how corporations are perfectly willing to launder copyrighted works via LLMs when it's profitable to do so. We have to perpetually pay them for their works and if we break their little software locks it's felony contempt of business model, but they get to train their AIs on our works and reproduce them infinitely and with total impunity without paying us a cent. It's that "rules for thee but not for me" nonsense that makes me reach such extreme logical conclusions that I feel empathy for terrorists.
- thesz 1y ago> copyright shouldn't even exist to begin with. You then get trade secretes and guilds. Hardly an improvement.
- matheusmoreira 1y agoSecrets? Just leak them, it only has to happen once. Guilds? Revoke their privileges and protections, and there's nothing they can do about it. Absolutely an improvement. Information wants to be free. Stop criminalizing it and people will find a way to free it. And once it's out there it's over, there is no containing it.
- wakawaka28 1y agoPeople want to be paid for their work. If you don't let them, they won't do the work. "Information" does not have a mind of its own. Even when the idea of a thing is "out there" there is a lot of grunt work and special stuff that needs to be implemented to get the best outcomes. Nobody owes you that work for free. Regardless of what GPL copers say, it is very hard to make money with software without enforcing some access restrictions and IP. Open source is great when it works, but it does not work for most things nor is it at the leading edge for most things.
- matheusmoreira 1y ago
- blibble 1y ago> it's already been ruled as fair use, multiple times most countries don't have a concept of fair use but they nearly all have copyright law
- quantummagic 1y agoThat fact in itself is a worse injustice than anything the LLM companies are doing. At the very least, it should be open to use in reporting, parody, and critique. Having no concept of such fair-use is oppressive and stifling.
- Hamuko 1y agoFair use is not the only way to allow critique and/or parody.
- quantummagic 1y agoWhat are you talking about? You can call it whatever you want, but it amounts to fair-use if you're allowed to use something for the purposes of critique and/or parody.
- Hamuko 1y agoThat's not how it works. Specific carveouts in copyright are not the same thing as fair use.
- quantummagic 1y agoYou're playing word games. The point is that any system that has no concept of fair-use, no allowance for reasonable usage of copyrighted works, except where explicitly granted by the copyright holder, is inherently unjust and stifling to free expression. How such allowances are specified in law, is irrelevant pedantry. What matters is that the allowances are afforded.
- thesz 1y agoOne can train model with copyrighted code as it is fair use, fair enough. Are there any rulings about use of code generated by model trained on copyrighted code? I believe distinction is clear.