8 ms·
The claim that's being allowed to proceed is under 17 USC 1202, which is about stripping metadata like the title and author. Not exactly "core copyright violati
by 0xcde4c3db 2y ago
The claim that's being allowed to proceed is under 17 USC 1202, which is about stripping metadata like the title and author. Not exactly "core copyright violation". Am I missing something?
- anamexis 2y agoI read the headline as the copyright violation claim being core to the lawsuit.
- H8crilA 2y agoThe plaintiffs focused on exactly this part - removal of metadata - probably because it's the most likely to hold in courts. One judge remarked on it pretty explicitly, saying that it's just a proxy topic for the real issue of the usage of copyrighted material in model training. I.e., it's some legalese trick, but "everyone knows" what's really at stake.
- 0xcde4c3db 2y agoYeah; I think that's essentially where the disconnect is rooted for me. It seems to me (a non-lawyer, to be clear) that it's damn hard to make the case for model training necessarily being meat-and-potatoes "infringement" as things are defined in Title 17 Chapter 1. I see it as firmly in the grey area between "a mere change of physical medium or deterministic mathematical transformation clearly isn't a defense against infringement on its own" and "giant toke come on, man, Terry Brooks was obviously just ripping off Tolkien". There might be a tension between what constitutes "substantial similarity" through analog and digital lenses, especially as the question pertains to those who actually distribute weights.
- kyledrake 2y agoI think you're at the heart of it, and you've humorously framed the grey area here and it's very weird. Sans a ruling that, for example, computers are too deterministic to be creative, copyright laws really seem to imply that LLM training is legal. Learning and then creating something new from what you learned isn't copyright infringement, so what's the legal argument here? A ruling declaring this copyright infringement is likely going to have crazy ripple effects going way beyond LLMs, something a good judge is going to be very mindful of. Ultimately, this is probably going to require congress to create new laws to codify this.
- mikae1 2y agoAccording to us law, is the Internet Archive a library? I know they received a DMCA excemption. If so, you could argue that your local library returns perfect copies of copyrighted works too. IMO it's somehow different from a business turning the results of their scraping into a profit machinery.
- kyledrake 2y agoMy understanding is that there is no concept of a library license and that you just say you're a library and therefore become one, and whether your claim survives is more a product of social cultural acceptance than actual legal structures but someone is welcome to correct me. The internet archive also scrapes the web for content, does not pay authors, the difference being that it spits out literal copies of the content it scraped, whereas an LLM fundamentally attempts to derive a new thing from the knowledge it obtains. I just can't figure out how to plug this into copyright law. It feels like a new thing.
- quectophoton 2y agoAlso, Google Translate, when used to translate web pages: > does not pay authors Check. > it spits out literal copies of the content it scraped Check. > attempts to derive a new thing from the knowledge it obtains. Check. * Is interactive: Check. * Can output text that sounds syntactically and grammatically correct, but a human can instantly say "that doesn't look right": Check. * Changing one word in a sentence affects words in a completely different sentence, because that changed the context: Check.
- mikae1 2y agoScraping with the intent to capitalize: no check.
- quectophoton 2y agoEven ignoring the fact that programmatic access to translation seems to require payment, or that its parent company is doing the scraping (similar to how one would use CommonCrawl instead of doing the scraping themselves), I am actually in favor of taking in to account the intent behind it. "Give and take", "equal exchange", however people want to put it. I don't mind if someone uses publicly-accessible content and ignores its copyright to make another thing, as long as their result is publicly-accessible and they're prepared to have their copyright ignored in return. If you not only use the result of someone else, but also their process, then be prepared to have your process publicly-accessible too, with its copyright ignored. And so on. That's why I don't mind "unofficial" translations or subtitles (both copyright violations as soon as they are distributed) appearing on multiple sites. That's why I respect open-source licenses of projects that respect them. That's why I pay for some open-source software even if I don't have to. That's why I give credit to artists even when I use an image that I didn't make myself as profile picture (either from the internet or because I paid for it). That's also why I don't mind anyone ignoring my copyright as long as it's on "equal" terms ("if you vendor my code and pass it off as yours, that's tacit approval for someone else doing the same thing to you" kind of thing ("someone else" because, at least for code, it won't be me)). I only gave very specific examples, but I hope I was able to explain what I mean. The thing that I don't like, is the highly asymmetrical situation we're in with generative AI: because the result (the trained model) is not publicly accessible like a significant part of the content it was trained on; they only release a very limited interface to it.
- Kon-Peki 2y agoViolations of 17 USC 1202 can be punished pretty severely. It's not about just money, either. If, during the trial, the judge thinks that OpenAI is going to be found to be in violation, he can order all of OpenAIs computer equipment be impounded. If OpenAI is found to be in violation, he can then order permanent destruction of the models and OpenAI would have to start over from scratch in a manner that doesn't violate the law. Whether you call that "core" or not, OpenAI cannot afford to lose these parts that are left of this lawsuit.
- zozbot234 2y ago> he can order all of OpenAIs computer equipment be impounded. Arrrrr matey, this is going to be fun.
- Kon-Peki 2y agoPeople have been complaining about the DMCA for 2+ decades now. I guess it's great if you are on the winning side. But boy does it suck to be on the losing side.
- immibis 2y agoAnd normal people can't get on the winning side. I'm trying to get Github to DMCA my own repositories, since it blocked my account and therefore I decided it no longer has the right to host them. Same with Stack Exchange. GitHub's ignored me so far, and Stack Exchange explicitly said no (then I sent them an even broader legal request under GDPR)
- CaptainFever 2y agoAlso, is there really any benefit to stripping author metadata? Was it basically a preprocessing step? It seems to me that it shouldn't really affect model quality all that much, is it? Also, in the amended complaint: > not to notify ChatGPT users when the responses they received were protected by journalists’ copyrights Wasn't it already quite clear that as long as the articles weren't replicated, it wasn't protected? Or is that still being fought in this case? In the decision: > I agree with Defendants. Plai ntiffs allege that ChatGPT has been trained on "a scrape of most of the internet, " Compl. , 29, which includes massive amounts of information from innumerable sources on almost any given subject. Plaintiffs have nowhere alleged that the information in their articles is copyrighted, nor could they do so . When a user inputs a question into ChatGPT, ChatGPT synthesizes the relevant information in its repository into an answer. Given the quantity of information contained in the repository, the likelihood that ChatGPT would output plagiarized content from one of Plaintiffs' articles seems remote. And while Plaintiffs provide third-party statistics indicating that an earlier version of ChatGPT generated responses containing signifi cant amounts of pl agiarized content, Compl. ~ 5, Plaintiffs have not plausibly alleged that there is a " substantial risk" that the current version of ChatGPT will generate a response plagiarizing one of Plaintiffs' articles.
- freejazz 2y ago>Also, is there really any benefit to stripping author metadata? Was it basically a preprocessing step? Have you read 1202? It's all about hiding your infringement.
- dragonwriter 2y ago”Core copyright violation”, here, I think is being used relative to the claims in the case.