3 ms·
Lawsuits based on what? Copyright? People crying for copyright in the context of AI training don't understand what copyright is, how it works and when it appli
by Lichtso 2y ago
Lawsuits based on what? Copyright?
People crying for copyright in the context of AI training don't understand what copyright is, how it works and when it applies.
What they think how copyright works: When you take someones work as inspiration then everything you produce form that counts as derivative work.
How copyright actually works: The input is irrelevant, only the output matters. Thus derivative work is what explicitly contains or resembles underlying work, no matter if it was actually based on that or it is just happenstance / coincidence.
Thus AI models are safe from copyright lawsuits as long as they filter out any output which comes too close to known material. Everything else is fine, even if the model was explicitly trained on commercial copyrighted material only.
In other words: The concept of intellectual property is completely broken and that is old news.
- LunaSea 2y agoLawsuits based on code licensing for example. Scraping websites containing source code which is distributed with specific licenses that OpenAI & co don't follow.
- Lichtso 2y agoUnfortunately not how it works, or at least not to the extend you wish it to be. One can train a model exclusively on source code from the linux kernel (GPL) and then generate a bunch of C programs or libraries from that. And they could publish them under MIT license as long as they don't reproduce any identifiable sections from the linux kernel. It does not matter where the model learned how to program.
- LunaSea 2y agoYou're mistaken. If I write code with a license that says that using this code for AI training is forbidden then OpenAI is directly going against this by scraping websites indiscriminately.
- Lichtso 2y agoSure, you can write all kinds of stuff in a license, but it is simply plain prose at that point. Not enforcable. There is a reason why it is generally advised to go with the established licenses and not invent your own, similarly to how you should not roll your own cryptography: Because it most likely won't work as intended. e.g. License: This comment is licensed under my custom L*a license. Any user with an username starting with "L" and ending in "a" is forbidden from reading my comment and producing replies based on what I have written. ... see?
- LunaSea 2y agoYou can absolutely write a license that contains the clauses I mentioned and it would be enforceable. Sorry, but the onus is on OpenAI to read the licenses not the creator. And throwing your hands in the air and saying "oh you can't do that in a license" is also of little use.
- CaptainFever 2y agoNo, it would not be enforceable. Your license can only give additional rights to users. It cannot restrict rights that users already have (e.g. fair use rights in the US, or AI training rights like in the EU or SG).
- LunaSea 2y agoHow does Fair Use consider commercial usage of the full content in the US?
- CaptainFever 2y agoIt's unknown yet, but the main point is that the inputs don't matter, as long as the output does not replicate the full content, it is fine.
- Lichtso 2y ago> You can absolutely write a license that contains the clauses I mentioned and it would be enforceable. A license (copyright law) is not a contract (contract law). Simply publishing something does not make the whole world enter into a contract with you. Others first have to explicitly agree to do so. > Sorry, but the onus is on OpenAI to read the licenses not the creator. They can ignore it because they never agreed to it in the first place. > And throwing your hands in the air and saying "oh you can't do that in a license" is also of little use. It is very useful to know what works and what does not. That way you don't trick yourself and your work to be safe, don't get caught by surprise if you are in fact not and can think of alternatives instead. BTW, a thing you can do (which CaptainFever mentioned) and lots of services do because licenses are so weak is to make people sign up with an account and have them enter a ToS agreement instead.
- jeremyjh 2y agoThat is not relevant to the comment you are responding to. Courts have been finding that scraping a website in violation of its terms of service is a liability, regardless of what you do with the content. We are not only talking about copyright.
- CaptainFever 2y agoTrue, but ToSes don't apply if you don't explicitly agree with it (e.g. by signing up for an account). So that's not relevant in the case of publicly available content.
- rcxdude 2y agoAlso, the desired interpretation of copyright will not stop the multi-billion-dollar AI companies, who have the resources to buy the rights to content at a scale no-one else does. In fact it will give them a gigantic moat, allowing them to extract even more value out of the rest of the economy, to the detriment of basically everyone else.
- lolc 2y agoAs much as our brain contents are unlicensed copies to the extent we can reproduce copyrighted work: If the model can recite copyrighted portions of text used in training, the model weights are a derivative work. Because the weights obviously must encode the original work. Just because lossy compression was applied the original work should still be considered present as long as it's recognizable. So the weights may not be published without license. Seems rather straightforward to me and I do wonder how Meta thinks they get around this. Now if the likes of Openai and Google keep the model weights private and just provide generated text, they can try to filter for derivative works, but I don't see a solution that doesn't leak. If a model can be coaxed into producing a derivative work that escapes the filter, then boom, unlicensed copy was provided. If I tell the model to mix two texts word by word, what filter could catch this? What if I tell the model to use a numerical encoding scheme? Or to translate into another language? For example assuming the model knows a bunch of NYT articles by heart, as was already demonstrated: If have it translate one of those articles to French for me, that's still an unlicensed copy! I can see how they will try to get these violations legalized like the DMCA safe-harbored things, but at the moment they are the ones generating the unlicensed versions and publishing them when prompted to do so.
- xdennis 2y ago> Lawsuits based on what? Copyright? > People crying for copyright in the context of AI training don't understand what copyright is, how it works and when it applies. People are complaining about what's happening, not with the exact wording of the law. What they are doing probably isn't illegal, but it _should_ be. The problem is that it's very difficult for people to pass new legislation because they don't have lobbyists the way corporations do.
- jcranmer 2y agoWith all due respect, the lawyers I've seen who commented on the issue do not agree with your assessment. The things that constitute potentially infringing copying are not clearly well-defined, and whether or not training an AI is on that list has of course not yet been considered by a court. But you can make cogent arguments either way, and I would not be prepared to bet on either outcome. Keep in mind also that, legally, copying data from disk to RAM is considered potentially infringing, which should give you a sense of the sort of banana-pants setup that copyright can entail. That said, if training is potentially infringing on copyright, it now seems pretty clear that a fair use defense is going to fail. The recent Warhol decision rather destroys any hope that it might be considered "transformative", while the fact that the AI companies are now licensing content for training use is a concession that the fourth and usually most important factor (market impact) weighs against fair use.
- Lichtso 2y agoLawyers commenting on this publicly will add their bias to reinforce the stances of their clientele. Thus somebody usually representing the copyright holders will say it is likely infringing and someone usually representing the AI companies will say it is unlikely. But you are right, we don't know until president is set by a court. I am only warning people that laying back and hoping that copyright will apply as they wish is not a good strategy to defend your work. One should consider alternative legal constructs or simply not releasing material to the general public anymore.