4 ms·
This is absurd. Remove all of the content from the training data that was pirated and what is the quality of the end product now?
by j_w 1y ago
This is absurd. Remove all of the content from the training data that was pirated and what is the quality of the end product now?
- pyman 1y agoWith Claude, people are paying Anthropic to access answers that are generated from pirated books, without the authors permission, credit, or compensation.
- KoolKat23 1y agoThere is no copyright on knowledge. If it outputs parts of the book verbatim then that's a different story.
- pyman 1y agoLet's don't change the focus of the debate. Pirating 7 million books, remixing their content, and using that to power Claude.ai is like counterfeiting 7 million branded products and selling them on your personal website. The original creators don't get credit or payment, and someone’s profiting off their work. All this happens while authors, many of them teachers, are left scratching their heads with four kids to feed
- KoolKat23 1y agoThat may be the case, but you'd have to have laws changed.
- SirMaster 1y ago>If it outputs parts of the book verbatim then that's a different story. But it does...
- KoolKat23 1y agoNot really, these have to be meaningful chunks and then yes, this is a separate court case.
- KoolKat23 1y agoThat's the law. Please keep in mind, copyright is intended as a compromise between benefit to society and to the individual. A thought experiment, students pirating textbooks and applying that knowledge later on in their work?
- j_w 1y agoWhen you say that's the law, as far as I'm aware a single ruling by a lower court has been issued which upholds that application. Hardly settled case law.
- KoolKat23 1y agoTrue, until then best to act as if it is the case. In my opinion, it will be upheld. Looking at what is stored and the manner which it is stored. It makes sense that it's fair use.
- j_w 1y agoWe're talking about a summary judgement issued that has not yet been appealed. That doesn't make it "settled." If by "what is stored and the manner which it is stored" is intended to signal model weights, I'm not sure what the argument is? The four factors of copyright in no way mention a storage medium for data, lossless or loss-y. (1) the purpose and character of the use, including whether such use is of a commercial nature or is for nonprofit educational purposes; (2) the nature of the copyrighted work; (3) the amount and substantiality of the portion used in relation to the copyrighted work as a whole; and (4) the effect of the use upon the potential market for or value of the copyrighted work. In my opinion, this will likely see a supreme court ruling by the end of the decade.
- KoolKat23 1y agoThe use is to train an AI model. A trillion parameter SOTA model is not substantially comprised of the one copyrighted piece. (If it was a Harry Potter model trained only on Harry Potter books this would be a different story). Embeddings are not copy paste. The last point about market impact would be where they make their argument but it's tenuous. It's not the primary use of AI models and built in prompts try to avoid this, so it shouldn't be commonplace unless you're jail breaking the model, most folk aren't.