5 ms·
I don't understand why it's even a question that Meta trained their LLM on copyrighted material. They say so in their paper! Quoting from their LLaMMa paper [To
by JW_00000 2y ago
I don't understand why it's even a question that Meta trained their LLM on copyrighted material. They say so in their paper! Quoting from their LLaMMa paper [Touvron et al., 2023]:
> We include two book corpora in our training dataset: the Gutenberg Project, [...], and the Books3 section of ThePile (Gao et al., 2020), a publicly available dataset for training large language models.
Following that reference:
> Books3 is a dataset of books derived from a copy of the contents of the Bibliotik private tracker made available by Shawn Presser (Presser, 2020).
(Presser, 2020) refers to https://twitter.com/theshawwn/status/1320282149329784833 https://twitter.com/theshawwn/status/1320282149329784833. (Which funnily refers to this DMCA policy: https://the-eye.eu/dmca.mp4 https://the-eye.eu/dmca.mp4)
Furthermore, they state they trained on GitHub, web pages, and ArXiv, which are all contain copyrighted content.
Surely the question is: is it legal to train and/or use and/or distribute an AI model (or its weights, or its outputs) that is trained using copyrighted material. That it was trained on copyrighted material is certain.
[Touvron et al., 2023] https://arxiv.org/pdf/2302.13971 https://arxiv.org/pdf/2302.13971
[Gao et al., 2020] https://arxiv.org/pdf/2101.00027 https://arxiv.org/pdf/2101.00027
- gameshot911 2y agoCritically, by torrenting they also directly distributed the copywritten material itself. That is a standalone infringement separate from any argument about trained LLMs.
- qup 2y agoAnd punishing them in the normal manner will be an incredibly small slap on the wrist, and do absolutely nothing to help us find out what will play out in court regarding a fair-use defense on training AI with copyrighted material.
- lucianbr 2y agoIsn't there a "fruit of the poisoned tree" kind of thing? Sounds to me quite similar to the situation where you would murder your parent and get to keep the inheritance, even if you are convicted of murder. Inheriting stuff isn't illegal, yet, I think most jurisdictions would not allow you to keep it in this case. There should be a problem with stuff obtained through illegal means, even if having that stuff is in principle legal. In this case, copyrighted material. Obviously they would argue that having the data is only a consequence of the download part, and that part is legal. What I see is that these situations are always complicated, and if you're rich enough, you get to litigate the complications and come out with a slap on the wrist or maybe even clean hands, while if you are an ordinary citizen, you can't afford to delve into the complexities and get punished. These days I'm starting to give up on the whole concept of the legal system being fair. They're not even pretending anymore.
- jimjimwii 2y agoThey could have only leached and refrained from sharing any part of copyrighted data. If i were to commit something as risky as this, that is what i would do.
- zelphirkalt 2y agoThen it would need to be determined, whether that is the case or not. Did every single machine they used have the configuration for only leeching and no seeding? The company is liable for what its employees on the job. If only one employee was also seeding ... that could be a very interesting case.
- crazygringo 2y ago> Did every single machine they used have the configuration for only leeching and no seeding? I would certainly assume so. It's incredibly obvious that's what you would want to do from a legal standpoint. > If only one employee was also seeding ... that could be a very interesting case. The torrenting wouldn't be done casually by employees acting on their own. And it's not like multiple employees are doing it simultaneously, unsupervised, on their personal computers. This is part of an official project. They'd spin up a machine just to download the torrent, being careful to disable seeding. This is Meta. They have lawyers involved and advising. This isn't a teenager who doesn't fully understand how torrenting works.
- mvdtnz 2y agoDid you not read the article? There are quotes from Meta employees doing exactly what you claim they wouldn't do. > This is part of an official project. They'd spin up a machine just to download the torrent, being careful to disable seeding. From the article: > "Torrenting from a corporate laptop doesn’t feel right," Nikolay Bashlykov, a Meta research engineer, wrote in an April 2023 message, adding a smiley emoji. In the same message, he expressed "concern about using Meta IP addresses 'to load through torrents pirate content.'" You also claim they would be "careful to disable seeding" but we know they did in fact seed (and anyone who uses private trackers knows they couldn't get away with leeching for very long before being kicked off): > Meta also allegedly modified settings "so that the smallest amount of seeding possible could occur," a Meta executive in charge of project management, Michael Clark, said in a deposition.
- Workaccount2 2y agoThere are two different things when it comes to discussing training LLM's on "copyright" protected data, and I almost never see people differentiate. 1.) Training on copyright that is publicly available. You write a poem and publish it online for the world to read. That is your IP, no one else can take it an sell it, but they are free to read and be inspired by it. The legalitly of training on this is in the courts, but so far seems to be going in favor of LLMs. 2.) Training on copyright that is not publicly available. These are pretty much pirated works or works obtained by backdoor to avoid paying for them. Your poem is behind a paywall and you never got paid, yet the poem is known by the LLM. This is just straight illegal, as you legally must pay to view the work. However there might be conditions here too like paying for access to an archive and then training on everything in it.
- farukozderim 2y agogood distinction IMO there's a hack about this, authors can claim that they allow for public use unless it's used for training LLMs. And all of training work would fall under 2 because they would be used against the copyright.
- echoangle 2y agoI think they would need to have some explicit contract every time they want to sell the book then, though. I don’t think I am bound by some random terms someone writes into a book I’m buying. Those probably are only binding if a reasonable person would notice them before sale.
- zelphirkalt 2y agoIf you arrive at the point of being able to buy that book, it means it has passed the publisher's hands and I would think, that the publisher was OK with those terms then, and limiting the usage of the text may in fact be effective. If it was self-published, then even more so.
- echoangle 2y ago
- unraveller 2y agoTrained on doesn't mean significant inclusion in the final state. Is it truly a violation of copyright when a user hacks out bits and pieces of easily restyled raw data points from a model to look samey? what about if it takes two models? Might be time to accept humans are just cooked in their ability to discern attempts at direct plagiarism - just as it is hard to discern Sky voice from Her voice.