5 ms·
I'm surprised the issue seems to be training on copyrighted material, that seems perfectly legal. I'm more interested in the ability of these models to violate
by aimor 3y ago
I'm surprised the issue seems to be training on copyrighted material, that seems perfectly legal. I'm more interested in the ability of these models to violate copyright by reproducing it. How much of a book is Llama2 allowed to generate before it's an issue?
- sillysaurusx 3y agoBingo. This seems to be the central issue. In fact, it’s not even copyright violation to distribute a model, since models aren’t copyrightable. They’re a collection of facts, like a phone book, produced by a purely automated process. It’s the outout that counts: https://news.ycombinator.com/item?id=36691050 https://news.ycombinator.com/item?id=36691050 And even the output has been ruled not copyrightable in recent court proceedings.
- sp332 3y agoSo you think training should be legal, but actually using the trained model to generate copyrighted text should be illegal? Edit: As a secondary point, it's not like Meta bought all those books. They downloaded unlicensed copies.
- harshreality 3y agoGoogle didn't buy the books they incorporated into books.google.com either. You can use books without buying them. e.g. libraries, or borrowing books from anyone even if they're not a "library". Copyright is not about internal use, it's about copying and distribution. This is not about what anyone thinks should be legal, but rather what is legal under the current law. The law was not designed for the digital era where "use" could be something other than a person consuming the content, and this has not been meaningfully addressed, therefore other uses are legal by default because that's how the law works.
- aimor 3y agoI'm outside of my expertise here, but I think if the model is able to reproduce copyrighted text then distributing the model should be a copyright violation.
- cdot2 3y agoA random text generator can generate copyrighted material.
- chii 3y agothe model can produce much more than just the original text. By your logic, distributing the digits of pi could be construed as copyright infringement otherwise.
- canjobear 3y agoSeems different if you can retrieve the text of a book by prompting the model with “Give me the text of X”. You can’t do that with the digits of pi.
- chii 3y ago> You can’t do that with the digits of pi of course you can - pi contains all known combinations of digits.
- sp332 3y agoThis is suspected, but has not actually been proven.
- canjobear 3y agoIf you wanted to extract a specific text from pi, you'd have to find the location of the text within pi. Expressed as an integer, this location would amount to an encoding of the entire text and would probably be longer than the text itself. You could only find the location by explicitly searching for the exact text. The location address would effectively be a copy of the text. On the other hand, the "address" of a text within a memorizing large language model would just be the prompt "give me the text of X".
- pixl97 3y agoAnd if you had 'read' them from a library?
- readyplayernull 3y agoI have a feeling this will lead to the creation of a system that compares works and the law will simply define a given threshold to conclude whether it's copyright violation or not.