7 ms·
“Copy” is ambiguous here. Of course data is copied during training. That said, OP is referring to whether the resulting model is able to produce verbatim copies
by thethimble 1y ago
“Copy” is ambiguous here. Of course data is copied during training. That said, OP is referring to whether the resulting model is able to produce verbatim copies of the data.
- triceratops 1y agoCopyright is the right to make copies. Why is copying during training is any different from producing copies of training data after training? If we're going that way, let me torrent every movie and TV show ever to "train" myself.
- foota 1y agoI don't think this is a reasonable argument. I don't think copyright is actually defined in that sense, but is perhaps more focused on consuming the content. Is an http proxy making a copy of something? What about computing an md5 of it as it's streamed through the proxy? Or maybe counting the words in the thing being served in order to track stats? I'd argue none of these fall under copyright, but each is an incremental step towards what it means to train a model.
- triceratops 1y ago> I don't think copyright is actually defined in that sense, but is perhaps more focused on consuming the content. https://en.wikipedia.org/wiki/American_Broadcasting_Cos.,_Inc._v._Aereo,_Inc https://en.wikipedia.org/wiki/American_Broadcasting_Cos.,_In.... I'm not a legal expert. My layman's understanding of the case above is Aereo was in violation because they made copies of content - content that the receiver was already allowed to access - available over the Internet to the intended receiver. That is to say, the copying was the problem.
- foota 1y agoI'm not a lawyer either, but from the summary a key part of the case seems to be that they were distributing the video to people for them to watch. "Aereo's retransmission of television broadcasts was a "public performance" of the networks' copyrighted work. The Copyright Act of 1976 forbids such performances without the permission of the holder of the copyright. Second Circuit Court of Appeals reversed. Court membership"
- atomicnumber3 1y agoIt's almost like information wants to be free
- deleted 1y ago[deleted]
- throwawaymaths 1y agoIt depends on your license? I mean strictly speaking if you stream a video you purchase legally over say amazon prime, there's lots of "copying" happening at various levels after those bits leave the data center.
- triceratops 1y ago> It depends on your license? Exactly this. Legal copying requires a license.
- ants_everywhere 1y agoYou can train yourself with every book at any library. You can also train yourself on a large number of movies and TV shows for a small monthly fee. Where your analogy goes wrong is you're saying you want to "[Circumvent] payment to obtain copyright material for training" to use Workaccount2's words.
- triceratops 1y agoSo Meta borrowed every book from a library and paid to obtain all of the movies and TV shows? They kept only one copy of every book at any time on their system? Because I'm certainly not allowed to photocopy a library book in its entirety. And I guarantee you a Netflix subscription doesn't allow me to keep a copy of a movie on my hard drive and use it for training man or machine.
- throwawaymaths 1y ago> Because I'm certainly not allowed to photocopy a library book in its entirety. IANAL but that probably falls under fair use? You'll get in trouble if you photocopy the work and sell access to it.
- 3036e4 1y agoDepends on where you live. In Sweden you can make a few copies of almost anything without violating copyright. There are a few exceptions. Copying entire books was added as an exception in 2005. You can still copy parts of a book. How large parts? I don't know, but I once asked for a copy from a library and they said that a few chapters was fine, so maybe that much (I am not a lawyer).
- triceratops 1y ago> but that probably falls under fair use? I've not found case law for that. I've had this same argument on HN multiple times over the past few months.
- cgriswald 1y agoCopyright is defined in law and as the original poster stated, whether this is 'copying' as defined by copyright law is legally ambiguous. Copyright doesn't protect against all forms of duplication. For instance, you own the copyright to your post and grant HN a license to offer copies of it. I have no direct license from you to copy the content of your post; but I can copy it to memory, copy a cache to disk, and make a copy appear on my display.
- kergonath 1y ago> For instance, you own the copyright to your post and grant HN a license to offer copies of it. It’s not a good example, because if you grant a license you give them the right to make copies. The problem is not when Meta got licenses, it’s when they did not.
- Dylan16807 1y agoThis line of conversation is not specific to the pirated books but is making claims about AI training in general.
- xyzzy_plugh 1y agoWhy does it have to be verbatim? Seriously, this I don't understand. If I produce a terrible shakycam recording of a film while sitting in a movie theater, it's not a verbatim copy, nor is it even necessarily representative of the original work -- muddied audio, audience sounds, cropped screen, backs of heads -- and yet it would be considered copyright infringement? How many times does one need to compress the JPEG before it's fair use? I'm legitimately curious what the test is here.
- ilikehurdles 1y agoIf you read a book and later understand its plot but can only explain it in your own words, did you copy it? The model isn’t storing the book.
- kxrm 1y agoThe model doesn't "understand its plot". So I am not sure this is a good analogy.
- ilikehurdles 1y agoTo what extent connections in a neural network are analogous to connections between neurons in your brain is open to interpretation and study, but the point of the analogy is that in neither case is a copy being made.
- jamiek88 1y agoYeah but a copy IS made. A human just reads. The machine copies the full text then compresses a lossy copy in its weights. You keep dodging that with tortuous analogies of a human learning. I’m sure all these ‘clever’ questions would be useful if this trial was about humans but it’s not.
- Workaccount2 1y agoModel training works roughly by feeding the model a text excerpt and then hiding the last word in the excerpt. The model is then asked to "guess" what the final word is. It will then move around it's weights until the guess sufficiently matches the actual token. Then the process repeats. The training material is used to play this guessing game to dial in it's weights. The training data is picked up, used as reference material for the game, and then discarded. It's hard to place this far from what humans do when reading, because both are using the information to mold their respective "brains" and both are doing an acquire, analyze, discard process. At no point is training data actually copied into the model itself, it's just run past the "eyes" of the model to play the training game.
- HWR_14 1y ago"Of course data is copied during training" is copying. As far as I know, the law is consistent that temporary copies are also covered by the copyright act, and that's how some analogous cases were resolved.
- shawabawa3 1y agoIf I buy a book I'm free to print as many copies as I want inside my house It becomes illegal if I try to distribute those copies So the question is, does distributing an AI that has been trained on Harry Potter count as distributing Harry Potter?
- triceratops 1y agoThis is not correct. It's only true that no one will go to the effort of prosecuting you for keeping photocopies of books in your home. But copyright law doesn't allow you to do it.
- Dylan16807 1y agoTemporary copies are in the scope of copyright law, yes. But also, you are allowed to make them. Or reading a book via a computer would be illegal.
- triceratops 1y ago> But also, you are allowed to make them. Not of physical media. You're allowed to make archival copies of digital media. > Or reading a book via a computer would be illegal No you purchased a license (or your library did, in the case of e-borrowing) to read the book on a computer. That makes it legal.
- Dylan16807 1y agoI am allowed to point a webcam at my physical book and read off the screen, even though that makes digital copies of all the text.
- OtherShrezzing 1y ago>That said, OP is referring to whether the resulting model is able to produce verbatim copies of the data. The NYTimes in 2023 was able to demonstrate that the models can reproduce entire articles verbatim[0] with minimal coercion. [0]https://nytco-assets.nytimes.com/2023/12/NYT_Complaint_Dec2023.pdf https://nytco-assets.nytimes.com/2023/12/NYT_Complaint_Dec20...
- WillPostForFood 1y agoPerhaps this is not evidence that the NY Times article was copied, but that what the NY TImes writes is highly predictable.
- OtherShrezzing 1y agoThe obvious test for this would be to have the models produce an article from before and after its cutoff date and see if the output is still verbatim. It would be a remarkable quirk of statistics that, if given all text on the internet except for the NYTimes back catalogue, a model would produce any NYT article.
- kergonath 1y agoThat copying is already a violation. At least it was when regular people weee on the receiving end of the lawsuits.
- superkuh 1y agoThe US Federal government operates with the rule that if human eyes don't look at it it doesn't count as a copy or looking at it. This allows them to unconstitutionally spy and log all people's telecommunications. Applying it here it seems pretty clear that corps are within the established bounds. As are any human persons that want to train an LLM this way.
- triceratops 1y agoSomeone engaged in large-scale unconstitutional spying does not give two fs about incidentally doing some copyright violations to achieve the spying. These are entirely orthogonal considerations.
- jayd16 1y agoSo if they could produce verbatim segments, that would be a violation? The technology is certainly there and these companies need to work backwards to prevent that.
- moomin 1y agoA good way of thinking about this is: consider the case where the data in question is illegal. Could you get into trouble for not only having access to it but also making copies of it? There’s plenty of case law there…
- lsaferite 1y agoI would argue that as an individual, real person, obtaining content without a license and personally consuming that content is significantly different than a corporation doing the same. My rational is that distribution of that content is (or should be) the primary offense. If I work for a company and they direct me to collect a bunch of content without a license and then I pass that to other members in my team to train a model, I've now distributed that content at the direction of my employer. That should be the offense the company is tried for. Using content to train an LLM is not copying the content. I'm ignoring the silly "but actually" arguments about the content being in RAM so it's "copying". It's using the content to generate a statistical model of token (word-ish) relationships and probabilities. If you write content that is so original in it's wording and I train an LLM against it, then there is certainly the possibility that the LLM could be provoked the recall the exact words you used. You'd have to set the parameters just right to make it happen and I think that proper training would drastically lower if not remove that possible scenario. But even if it doesn't, the LLM doesn't have a copy of that original content. All it has is weights representing those relationship probabilities. Yes, the minutia is more complex, but that is the essence. If my LLM were to generate enough of this essentially verbatim unique content and I tried to publish or copyright it, then I as the user should be on the hook. But then you get into a discussion about how many words in a unique sequence does it take to be infringement? Obviously, I am not a lawyer. My summation in all of this is that new laws need to be put into place to handle this stuff because the existing ones are sufficiently non-definitive and/or ill-suited such that every party is forming strong opinions about how old laws apply to new situations and causing massive friction.
- crystal_revenge 1y ago> That said, OP is referring to whether the resulting model is able to produce verbatim copies of the data. Transformers are fundamentally large compression algorithms where the target of compression is not just to minimize reconstruction loss + compressed file size. In fact, basically all of machine learning used today can be viewed through the lens of learning a compression algorithm with added goals other than the usual. By this logic if I create a lossy Jpeg of a copyrighted image it's not "copying" because the lossy compression.
- dragonwriter 1y ago> That said, OP is referring to whether the resulting model is able to produce verbatim copies of the data. While a tool being used to create infringing copies of some other work (whether or not it is the source material used to create the tool, and whether or not the infringing material is also verbatim copies) is relevant to whether the tool vendor is liable for contributory infringement for the infringing use of the tool, the absence of a capacity for creating such copies isn't usually enough to say that copying to make the tool isn't infringing. (That said, generative AI tools, including LLMs specifically, have been shown to have the capacity to make such copies, to the extent that vendors of hosted models are now putting additional checks on output to try to mitigate the frequency with which verbatim copies of substantial portions of training-set works are produced, so arguing that LLMs can't do that is silly.)
- arh68 1y ago> LLMs specifically, have been shown to have the capacity to make such copies Exactly. I asked my Gemma how long of a quote it could give me of a given book, if I were the author & gave express permission, and I was a bit surprised it readily admitted it could > Without Permission (Current Limit): Single sentence. > With Broad Permission (Full Reproduction Allowed): I could theoretically quote the entire book. Eye-opening (for me, at least).