7 ms·
I don't think the MP3.com case is a good comparison, because copyright law is adjudicated on the output being infringing material, and in the case of MP3.com, t
by carlosdp 3y ago
I don't think the MP3.com case is a good comparison, because copyright law is adjudicated on the output being infringing material, and in the case of MP3.com, the output was the exact songs they were storing. That's a much more clear cut case.
In the case of OpenAI et al, the NYT is claiming that because they were used in inputs, the whole model is infringing. Or, because the model can be coerced into producing a copyrighted output, the whole thing is infringing.
But I agree with Jeff, that's now how a judge will see it. Individual outputs from a model can be claimed as infringement (for example, the Mario reproductions) against any user that publishes those works, but that does not make the models themselves infringing simply because they observed copyrighted works while training. That's clearly fair use.
- philipodonnell 3y agoIs OpenAI as a company more like a publisher or a model?
- polski-g 3y agoOpenAI is more like a subscriber to the Times. Somebody who reads an article about housing policy, and writes their own paragraph on Facebook utilizing ideas from the Times article.
- tpmoney 3y agoPersonally I think OpenAI is more like Xerox. They’ve invented this device (their AI models) that can be used with the right inputs to generate copyright infringing outputs. But it still requires a user to generate those outputs by choosing the inputs. On its own it’s just a tool that’s no more copyright infringing than any other photocopier.
- asadotzler 3y agoI don't agree. Xerox didn't require stealing the worlds books in order to make copies. Those machines don't have every book inside and spit out the pages you request, they're just cameras, basically, taking pictures of what's put in front of them. That's very different from a tool that must first steal all the content in the world before it makes any outputs.
- tpmoney 3y ago1) OpenAI didn't steal anything. 2) The models explicitly don't have every book inside. At best (which the court case is dealing with in part) one could argue that OpenAI is transforming works covered under copyright in a manner that isn't sufficiently transformative to pass the fair use exceptions to copyright and is thus committing copyright infringement. It's still not theft. OpenAI doesn't "require stealing the world's books" in order to do what they do. Their product is vastly more effective and useful because it was trained on such a wide corpus of material, but likewise a xerox machine that won't make any copies of anything under copyright is vastly less useful and effective as one that will. Likewise a VCR that refuses to record from TV is less useful than one that will. A CD drive that refuses to rip MP3s from CDs is less useful than one that will. A BitTorrent client that refuses to send or receive items subject to copyright is less useful than a client that will send and receive those items. The fact that a product is better by it's ability to be usable in committing copyright infringement is neither evidence that the product itself is infringement, nor a strong argument in favor of preventing the product from being capable of such infringement.
- asadotzler 3y agoYou must have missed the NYT's suit and initial evidence. Not a problem, you'll fit right in with all the other temporarily embarrassed millionaires simping for this plundering of the commons..
- Kim_Bruning 3y agoAlso, if the model doesn't know what Mario looks like, you can't give it a negative prompt to NOT produce Mario.
- feoren 3y ago> because they were used in inputs, the whole model is infringing That is a nonsense argument on its face, but: > because the model can be coerced into producing a copyrighted output That is an extremely good argument. In fact it's almost completely damning, unless you can show that "being coerced into it" means re-introducing enough information in the input that the model's infringement is mostly a regurgitation of that input. But it's clearly not, if you can simply ask it nicely to reproduce an entire text and it does so. > Individual outputs from a model can be claimed as infringement (for example, the Mario reproductions) against any user that publishes those works How is that different than operating MP3.com, or a Warez site? The same argument would say "it's the downloading user that is infringing, not the platform!" But clearly that hasn't held up. Consider ChatGPT as a platform hosting a ton of copyrighted material that it produces for you if you ask it nicely, and it's clearly in a much worse position than MP3.com was. ChatGPT is itself publishing those works. Even if it's not hosted online, publshing the GPT model means publishing a huge collection of copyrighted works. I don't see any way around it.
- kromem 3y agoThe model producing copyrighted material isn't as great an argument as people seem to think. The cases are pursuing training as infringement, not usage. So in the case of Mario - there's no infringing in learning the attributes of the most recognizable Italian plumber in video games. It is only when the models create images of Mario in their usage that's infringing (which will be separate cases and those will likely be a shoe-in for plaintiffs forcing copyright tagging filters in front of publicly accessible generative models). The most damning part is the reproductions of the NYT text, as that's not simply learning attributes of a copyrighted character, but verbatim partial duplication. I suspect in many of those cases it's due to fair use copying of segments of NYT articles by multiple other sources in the training data, but it's going to be difficult waters for the defense to navigate even if that's the case - but this is also technically impossible for any trained LLM to avoid. If a source you have rights to quotes a source you don't have rights to, you are going to ingest a legally permissible usage of material that suddenly will no longer be legally permissible to have ingested? We'll see how it plays out, but it really seems like it's just going to come down to a drawn out appeals battle no matter how it lands given different judges are likely going to each ultimately see it differently given both sides have potentially compelling arguments.