3 ms·
Fair use allows for certain uses of copyrighted works without a specific license for those works. One of the major criterion is how transformative the work is
by TimPC 1y ago
Fair use allows for certain uses of copyrighted works without a specific license for those works. One of the major criterion is how transformative the work is and an LLM model is very different from the original work so it seems likely that criterion at least is met.
- SideburnsOfDoom 1y ago> an LLM model is very different from the original work True, but not the only relevant thing. If the output of the LLM is "not very different from the original work" then the output could be the infringement. Putting a hypercomplex black box between the source work and the plagiarised output does not in itself make it "not infringing". The "LLM output as a service" business is then based on selling something based other people's work, that they do not have rights to. It's falling for misdirection, "pay no attention to the LLM behind the curtain" to think otherwise.
- Filligree 1y agoThe output of the LLM is very different from the original, though. It’s hard to look at this and claim it isn’t.
- SideburnsOfDoom 1y ago> The output of the LLM is very different from the original, though I will disagree with that characterisation. IMHO: In some cases no, it's not different, there are clear lines from inputs to output. In some cases yes, it's different from any one input work, it's distributed micro-plagiarism of a huge number of sources. In no case is it original. But I think that this is legally undecided and won't be decided by you or me, and it is going to be a more interesting and relevant question than "is the LLM model is very like the original work", which it clearly isn't. That's like asking "is this typewriter like this novel?" It can't be, but the words that came out of it could be.
- ijk 1y agoYeah, the way the courts decide is unlikely to turn on a detail of how the technology works, so it's difficult for non-legal experts to predict the outcome on the technical merits (since the law has very different priorities). Music has ended up in a place where short audio snippets are protected by copyright and must be licensed; but for short snippets of text the precedent has generally been that the copying needs to be more substantial. Distributed microplagarism of short phases might end up being ruled to be legal, even if wholesale reproduction is not. Which may not give copyright protection to the generated works, of course, as the question of machine authoring is entirely distinct.
- sillysaurusx 1y ago> In some cases yes, it's different from any one input work, it's distributed micro-plagiarism of a huge number of sources. In no case is it original. That’s like saying the dictionary is micro-plagiarism of a huge number of sources because it uses all the words from those sources.
- SideburnsOfDoom 1y agoI disagree, you can't ask a dictionary to "generate 2000 words in the style of (author)".
- sillysaurusx 1y agoSo? Why is that, of all things, the crux of whether it’s copyright infringement? Plagiarism isn’t necessarily copyright infringement, and plagiarism isn’t illegal. Copyright infringement is. Even still, your argument that everyone who generates 2,000 words in the style of (author) is plagiarizing is also flatly false. By that standard all English essays that mimic someone else’s style would be plagiarism.
- int_19h 1y agoIn the general case, yes, but they can verifiably reproduce at least some copyrighted works verbatim, which implies, at the minimum, that their content is stored in model weights in some fashion.
- lsaferite 1y agoIt implies that the token procession probability was unique enough that with a low entropy token stream and the proper starting token stream you could recreate portions of the original content *strictly based on probabilities*.
- johanyc 1y agoEveryone knows the training data is stored in some way in the LLM. The point is the use of the copyrighted material is transformative. Remember google books, it literally shows photocopy of pages of books but the court ruled it’s fair use. A simplified explanation is book vs search engine and book vs ai chatbot are very different from each other.
- SideburnsOfDoom 1y ago"a photocopy of pages of books" is exactly that, pages of an existing book. It doesn't pretend to be something else. The output of a LLM, when based heavily on that same page, pretends to be something novel. IMadeThis_Meme.gif