8 ms·
I raised a concern about the inclusion of books3 in the Pile back in 2020, and this is what the head of Eleuther (Stella Biderman) told me: "So here’s the big
by Ninjinka 3y ago
I raised a concern about the inclusion of books3 in the Pile back in 2020, and this is what the head of Eleuther (Stella Biderman) told me:
"So here’s the big picture. There are three sets of datasets:
1. Data exists out there in the world. It has been collected into datasets and posted online. I’ll call this raw data.
2. We take that data, clean it, and process it for language modeling. I’ll call this per-set data.
3. We combine those per-set data into one massive dataset, the Pile. This is heavily processed, including weighing the components.
We created 2 and 3 and put them online. We put 2 online so that people can reweigh and remix the data if they wish, but we expect most people to just download 3 and use it out of the box. Access to 3 will be provided in several forms, including HuggingFace and from our website.
2 and 3 are not copyright violations, even if the data is copyrighted, because they fall under fair use (at least in the US).
The Pile contains code that turns 1 into 2 and code that turns 2 into 3.
When you download Maroon 5 from a website, you are creating a dataset corresponding to 2. That can be copyright violation depending on what you do with it, but our use is not a copyright violation."
- anticensor 3y agoIn Europe, 2 and 3 are subject to compilation copyright and database rights.
- artninja1988 3y agoHopefully that is correct. The pile has been very valuable for open model work. It's a really high quality dataset
- D-Se 3y ago[flagged]
- layer8 3y agoI don’t understand how this can be true if set 2 contains a complete copyrighted work (say, a book) that the copyright owner hasn’t approved for such distribution. Unless I misunderstand and the “process[ing] for language modeling” is an entirely irreversible process.
- doctorpangloss 3y ago[flagged]
- layer8 3y agoYour last sentence is a non-sequitur.
- doctorpangloss 3y agoAll of this conversation aside, anyone who is saying "fair use" concedes that their use more or less would be explicitly unauthorized, but because it advances "some" "goals" in this country, it should be allowed anyway. It is a necessarily adversarial stance but it belies that the author the Pile sincerely wants to advance human progress in an idiosyncratic way. I think this is what pisses off the commenters here, they want to take sides, but like, who cares. > I don’t understand how this can be true if set 2 contains a complete copyrighted work (say, a book) that the copyright owner hasn’t approved for such distribution. Everything you're saying could be true, but also meaningless. Meaningless in what sense? In my personal opinion, here are some meaningful things: - Economic meaning: Is the Pile making anyone suffer, in an intellectually honest sense? No. Is the way it is being used making anyone suffer? No. Would anyone use the Pile if they had to pay for it? No. Would some tiny amount of money paid for it matter? No, not to anyone. The Pile is but one of many things that are going on in this pareto-inefficient world that we gain little economically by fucking with. - Meaning in the sense of human progress. Is the Pile sincerely, substantively helping advance creative and scientific progress? Yes. Can the progress it promotes coexist peacefully, in this current status quo, with selling books on Amazon, which can also advance progress? Yes. Is there any evidence that this non-economic progress, such as sharing important stories or points of view meaningfully through text content, is harmed by the Pile? No. The following ways of looking at this issue are meaningless: - Psychological meaning is meaningless. Feelings that are more strongly felt by more favored parties, like beloved authors versus impetuous programmers, are not more valid. Stakeholders like authors merely exercising their collective bargaining power over something that doesn't actually make them worse off is meaningless too. Social media drama also doesn't matter. These are just opinions. I don't go out there and say I'm a lawyer, and I'm not running for Congress, and neither are you. It's crazy to me, because when you look critically at what dog you personally have in this race, like of course you want the status quo where The Pile exists, and you and the only agitators in all of this - authors' guild authors - gain nothing from nearly every dead-on-arrival framework like "Compensation, Credit and Consent" or whatever, BESIDES meaningless psychological satisfaction of winning social media arguments and flexing the favorability of authors over programmers. Imagine if we ran the whole country this way! It's radioactive. So this isn't a non-sequiter. Supporting the status quo, in this particular case, is the succinct way of saying everything here: that you can be right in ways that don't at all matter, which is okay, but you should have the insight or maybe if you are a programmer or an author or actually you are running for Congress, you have the duty to understand this stuff.
- camkego 3y ago[flagged]
- chinathrow 3y agoNicely stated copyright violations. Has noone filed suit yet?
- SEGyges 3y agoHuckabee v Bloomberg, Meta, et al
- whimsicalism 3y agoscraping libgen and downloading copyrighted content and redistributing it isn’t illegal? call me skeptical, seeding a torrent of movies that you downloaded from elsewhere on the internet isn’t “fair use” and the pile isn’t just code for transforming data, it is the redistributed data itself by this logic i could legally run a libgen mirror
- nickpsecurity 3y agoThey’re distributing copyrighted works without the authors permission, using them in ways that compete with the author, many make money off AI’s, and the AI’s reproduce some verbatim. These datasets seem to fail most tests ("four factors") in copyright law. Even laypeople I’ve explained LLM’s to think the AI companies are ripping others’ work off. For those concerned, I have an article that covers legalities, each dataset (including The Pile), legal issues with them, alternatives that are legal, and a copyright amendment that balances all sides. http://gethisword.com/tech/exploringai/ http://gethisword.com/tech/exploringai/ Looking back at my proposal, I think we need at least three rules passed immediately in at least one country: 1. All copyrighted works can, if a person has legal access, be used for training AI systems. Any terms restricting copyrighted works from use in training, charging more for that, restricting downloads for it, etc are illegal. Every act of publishing can benefit both a human mind and AI training equally. 2. People can copy and transform for their own use any work they have access to only for AI training. This might include reverse engineering for extraction, multiple copies in different formats, and so on. They can do whatever is needed to get it into the AI system. Other uses or abuse of this data is subject to existing law. 3. Any work published online for free and with public access can be copied, shared, processed, and bundled for AI training. That’s regardless of its terms. Note: In No. 2 and No. 3, the resulting AI’s copyright will be determined by existing law about AI’s and mixing copyrighted works. Or no copyright if that’s the law. 4. If AI outputs are copywritten, their status will be the same as if the user published it themselves while relying on prior works. AI training sets will also be public to determine this. With those rules, we can share works like those in The Pile, still pay creators that want to be paid, be less likely to just steal existing work, and infringement in outputs is still illegal. What do you all think of that?
- dougb5 3y agoI don't know what the right answer is to the copyright questions, but I hope that in 2024 we'll have a better attitude about the human labor that went into these models than "Data exists out there in the world" and the passive-voice "It has been collected into datasets"
- otterley 3y ago> 2 and 3 are not copyright violations, even if the data is copyrighted, because they fall under fair use (at least in the US). This cannot be known until it is litigated. Fair Use is not something you can unilaterally declare and have it be so, just like you can't be like Michael Scott in the Office shouting "I declare bankruptcy!" OpenAI is currently defending itself against the New York Times for this very reason. There's a multi-factor test that courts weigh the facts against in making a determination as to whether a prima facie copyright violation would be protected under a Fair Use defense: Factor 1: The Purpose and Character of the Use Factor 2: The Nature of the Copyrighted Work Factor 3: The Amount or Substantiality of the Portion Used Factor 4: The Effect of the Use on the Potential Market for or Value of the Work See https://copyright.columbia.edu/basics/fair-use.html https://copyright.columbia.edu/basics/fair-use.html for a pretty good overview of what the analysis entails.
- ryukoposting 3y agoThanks. This is really informative, and really important information given the growing relevance of IP law in everyone's daily life. Part of me wonders if these four factors will ever become part of core curriculum for civics classes. By no means am I an expert in copyright law, but factor 3 seems like very bad news if you're OpenAI.
- fluoridation 3y agoWhether something is fair use or not is not determined by a court, but by the definition of what fair use is. A court interprets that definition and the situation, and if their interpretation matches yours you may have a ruling in your favor. But saying "this is fair use" is no more incorrect than saying "this is red". You're interpreting your perception and putting that interpretation into words.
- deleted 3y ago[deleted]
- otterley 3y ago> But saying "this is fair use" is no more incorrect than saying "this is red". When a court determines that it isn't, you can continue to argue it as much as you like (to deaf ears), and yet you're still liable to the copyright holder. Whether it's "incorrect" or not is then irrelevant. Let's not argue semantics here.
- 3abiton 3y agoInteresting take on the copyright law.
- tycho-newman 3y agoFair use is a defense to infringement. Do not start your copyright argument by admitting you infringed.
- MacsHeadroom 3y agoFair use is an exception to infringement. Use which is fair is non-infringing. For example, Google books containing a searchable copy of every book ever written is fair use, as is Google's cache containing every news article and web page.