4 ms·
Copying verbatim passages can be fair use if it is transformative.
by mannerheim 3y ago
Copying verbatim passages can be fair use if it is transformative.
- tyingq 3y agoOutputting unattributed copies large passages of text to end customers with the implied okay to use it any way they want, though....that's what these tools do. The end user often has no idea they just received something with potential IP issues.
- lsaferite 3y agoDoes the model even have enough information to be able to _know_ that though? If it's simply using probability to ties a series of tokens together into text based on numerical probability, that doesn't store enough information to understand that _this specific sequence of tokens_ represents come specific copyrighted work. Storing that information would _actually_ seem to fall afoul of copyright. The fact that it can spit out chunks of copyrighted works is driven by the input token sequence and the model weights pointing to a specific path that have an ever so slightly higher probability of being the expected output, right? It's not like the model stores the copyrighted work directly. (Yes, I know the algorithms are more complex than what I expressed, but the general idea holds in my understanding)
- pclmulqdq 3y agoThe model doesn't know anything - people personify LLMs too much. It is a mathematical text predictor that has almost certainly ingested the text it is copying verbatim to string together the words it is reproducing. The fact that it is a highly compressed representation of its training corpus (and thus doesn't "know" that it is copying something) is not an excuse. I think you could make a good argument about this if you could prove that the text being spit out verbatim is _not_ contained in the training corpus, but that is not the situation we have today.
- lsaferite 3y ago> The model doesn't know anything - people personify LLMs too much. Perhaps reading people's posts little less literal would help the conversation. I obviously know that the model weights don't 'know' something in the sense that a human knows something, but the model does store information. That information happens to (primarily?) be the statistical likelihood of one token following another token. What it doesn't store is a string of tokens that represents Sarah Silverman's latest book. From what I can tell, all this angst comes from 3 or 4 related, but different, issues. 1. Did companies break copyright laws when assembling and using training for these models? 2. Does the model represent some form of copyright infringement in and of itself? 3. Does a model's ability to output chunks of copyrighted work have some implication of the legality of the model itself? (using said copyrighted chunks is already a solved issue) 4. Do we as a society owe it to humans benefitting from copyright the continued ability to create copyrightable content without competition from ML models? I think comingling all of those points is doing everyone a disservice. My assertion was only about #2 and none of the others. I feel like it's a clearly demonstrable situation that these models _don't_ infringe copyright directly. That being said, I am obviously not a lawyer, and my opinion is just that, an opinion. FWIW, my general feeling on all of the points is: 1. Quite Likely (but fair-use is a fickle thing), 2. No, 3. It shouldn't, and 4. No, but we need to think through the long-term societal implications of ML decreasing the amount of human labor needed across all markets and come up with a plan that doesn't involve our fingers in our ears.
- pclmulqdq 3y agoI would assume that #2 is actually "yes" given the way derived works work and the fair use tests, and that #1 and #2 are actually very much linked. Fair use is a defense to copyright infringement, and it's a relatively complex balance of factors. It's relatively inarguable (even OpenAI isn't arguing this) that the model isn't a derivative work of its training set: they are just arguing that they have fair use rights to the contents of the training set. Only one of the factors in a fair use analysis is how transformative the use is, and I think it's hard to argue that training an LLM isn't a huge transformation. However, the other factors weigh pretty heavily against LLMs here, and the Author's Guild lawsuit is a pretty good set of arguments as to why. It's up to a court to decide whether the transformative nature outweighs the other factors. If you're lumping the fair use question into #1 and the "is it a derived work" question is #2, I'm pretty sure that nobody on any side of this agrees that it isn't a derived work. Once you have a derived work, you can either ask whether they were licensed to produce that work (no) or whether it was fair use (possibly).