29 ms·
So then it needs to be proved that real people actually cancelled their NYT subscription in favor of reading articles from chatGPT. This is incredibly unlikely
by MeImCounting 3y ago
So then it needs to be proved that real people actually cancelled their NYT subscription in favor of reading articles from chatGPT. This is incredibly unlikely for several reasons least of which is GPTs training cutoff, the need to copy past several paragraphs of the article to get GPT to finish it among other reasons.
- dahart 3y agoFortunately, that’s not how copyright works. Burden of proof is on OpenAI to justify their copying without permission, they don’t get to copy anything and then demand that anyone who doesn’t like it prove that financial damage was done first before they stop. Financial damage does count under copyright, but is not required to win an infringement claim. Courts or the copyright office might decide that AI can count as fair use, or they might continue to allow suits every time an AI spits out verbatim copyrighted work in violation of the law, or we might get some entirely new criteria for copyrights. It’s going to be interesting.
- MeImCounting 3y agoTo be honest I really dont understand how copyright works. I have read through the first several pages of the NYT case PDF but it still didnt make much sense to me. Is the issue just that GPT is able to repeat the article word for word or is the issue about the article being in the training data at all? If its not about losing business what is it about? Could you or someone else expand a bit on what exactly the problem is and maybe even speculate as to potential outcomes?
- dahart 3y agoThis is a good question, and I suspect the answer is some of both, but the lawsuit is very specifically claiming that ChatGPT is “memorizing” and reproducing NYT articles verbatim, and this is clearly in violation of the law. I suspect they’re also trying to push on the issue of it being illegal to train AI on the NYT archive, but that’s a bigger and more difficult question. The issue of verbatim reproduction makes a clear legal case, which is why they’re relying on that first, and letting that lead to the broader discussion. Exhibit J in the suit is titled: “One hundred examples of GPT-4 memorizing content from the New York Times”.