4 ms·
This feels like an extremely weak defense to me, because if it works it is essentially a copyright backdoor. 1. There is little doubt that OpenAI has NYT artic
by TrainedMonkey 2y ago
This feels like an extremely weak defense to me, because if it works it is essentially a copyright backdoor.
1. There is little doubt that OpenAI has NYT articles in it's dataset.
2. They claim it's OK because all of the IP is mixed together and effectively laundered.
3. But in fact if you put in a secret code you can get the IP back nearly verbatim.
4. So if OpenAI wins this case we could just trade prompts that regurgitate the articles back without ever visiting NYT.
- jrockway 2y agoI agree. If you gzip a NYT article and print it out, very few people would be able to read the article. But it can still be decoded ("prompt engineering" as OpenAI calls it).
- sidewndr46 2y agoCopyright maximalism in the 21st century can be summed up as: When an individual makes a single copy of a song and gives it to a friend, that's piracy. When a corporation makes subtlety different copies of thousands of works and sells them to customers, that's just fair use
- JumpCrisscross 2y agoNYT is a corporation. Corporation vs individual is a distraction. It’s some people (wrongly, in my view) prioritising production over consumption. If this were Altman personally producing an AI, the same people would rally to him. The corporate/individual framing needlessly inflames the debate when it’s really one about power and money.
- delusional 2y agoI don't think it's "production over consumption". At least I don't like that framing. For me it's about supporting production. The humans that write news articles every day can't produce that valuable work if they don't get fairly compensated for it. It's not that the AI produces more, it's that the AI destabilizes production. It makes it impossible to produce.
- JumpCrisscross 2y ago> It's not that the AI produces more We're not debating whether they do. "Humans that write news articles" are producing. That contrasts with "an individual mak[ing] a single copy of a song and giv[ing] it to a friend." We don't put journalists in jail for plagiarism.
- delusional 2y ago> We don't put journalists in jail for plagiarism. I'm guessing you're imagining a scenario here were a journalist has copied an entire article verbatim and republished it in their newspaper. That would actually be both copyright infringement AND plagiarism. Newspapers just rarely enforce that right. These two things aren't on a scale. They are independent infractions.
- freejazz 2y agoNo, they wouldn't, because Altman would still be stealing other people's actual work.
- JumpCrisscross 2y ago> they wouldn't, because Altman would still be stealing other people's actual work OpenAI is "stealing other people's [sic] actual work." The people rallying to it clearly don't care that much about it now. They wouldn't care whether it's a corporation or Sam Altman per se doing it.
- Maxatar 2y ago1. Anyone can get all of NYT's articles for free along with CNN and every other major news site, this isn't in dispute, it's available here in a single 93 terabyte compressed file: https://data.commoncrawl.org/crawl-data/CC-MAIN-2025-05/index.html https://data.commoncrawl.org/crawl-data/CC-MAIN-2025-05/inde... 2. I did not see any defense of this nature. 3. Yes and this is the big deal. If the secret code needed to reproduce copyrighted material involves large portions of that copyrighted material already then that's quite a bit different than just verbatim reproductions out of thin air. 4. Yes, if OpenAI wins this case then you could feed into ChatGPT large portions of NYT articles and OpenAI could possibly respond by regurgitating similar such portions of NYT articles in response.
- DSMan195276 2y agoI'd say there's some merit to that defense. Imagine for example if a website generated itself based on a sequence in Pi - technically all of the NYT is in that 'dataset' and if you tell it to start at the right digit it will spit back any NYT article. In a more realistic sense though you can make it spit back anything you want and the NYT article is just a consequence of that behavior - finding the right 'secret code' to get a verbatim article is not something you can easily just do. ChatGPT is somewhere in-between - You can't just ask it for a specific NYT article and have it spit it back at you verbatim (NYT acknowledges as such, it took them ~10k prompts to do it), but with enough hints and guesses you can coax it into producing one (along with pretty much anything else you want). The question then becomes whether that's closer to the Pi example (ChatGPT is basically just spitting the prompt back at you), or if it's easy enough to do that it similar to ChatGPT just hosting the article. Edit: I suppose I'd add, this is also a separate question from the training, training on copyrighted material may or may not be legal regardless of whether the model can verbatim spit the training material back out.
- throwway120385 2y agoI think the difference here is that a human intentionally built a dataset containing that information, whereas Pi is an irrational number which is a consequence of our mathematics and number system and wasn't intentionally crafted to give you NYT articles.
- DSMan195276 2y agoWell that depends on what you're trying to prove. If you think it's a copyright violation to include the articles in the dataset _at all_ then it doesn't even matter if ChatGPT can produce NYT articles, it's a violation either way. If including the articles in the dataset is not in-and-of-itself a copyright violation then things get complicated when talking about what prompt is required to produce a copyright-violating result.
- delusional 2y agoYou're getting lost in the technology here. Copyright is not about producing the exact sequence of bytes, nor is it about "hosting an article". Copyright is an intellectual property right to the creative work, not the exact reproduction that is seen on some website, but the creative work itself. The law doesn't not care about your weird edge cases. What matters is what should be and how we can make it so.
- AnthonyMouse 2y ago> So if OpenAI wins this case we could just trade prompts that regurgitate the articles back without ever visiting NYT. This seems like the inverse of the old "book cipher" scheme to "avoid" copyright infringement. If you want to distribute something you're not allowed to, first you find some public data (e.g. a public domain book), then you xor it against the thing you want to distribute. The result is gibberish. Then you distribute the gibberish and the name of the book to use as a key and anyone can use them to recover the original. The "theory" is that neither the gibberish nor the public domain book can be used to recover the original work alone, so neither is infringing by itself, and any given party is only distributing one of them. Obviously this doesn't work and the person distributing the gibberish rather than the public domain book is going to end up in court. So then which side of the fence is ChatGPT and which side is the text you have to feed it to get it to emit the article? Well, it's the latter that you need access to both the existing ChatGPT and the original article in order to produce. Notice also that this fails in the same way. The people distributing the text that can be combined with the LLM to reproduce the article are the ones with the clear intention to infringe the copyright. Moreover, you can't produce the prompt that would get ChatGPT to do that unless you already have access to the article, so people without a subscription can't use ChatGPT that way. And, rather importantly, the scheme is completely vacuous. If you already have access to the article needed to generate the relevant prompt and you want to distribute it to someone else, you don't have to give them some prompt they can feed to ChatGPT, you can just give them the text of the article.