4 ms·
So.. the data set is licensed under Creative Commons, the source images all have their copyright (so how did they make it into the data set?) but what about ima
by optymizer 4y ago
So.. the data set is licensed under Creative Commons, the source images all have their copyright (so how did they make it into the data set?) but what about images created with Stable Diffusion that uses this data set? Are those derived works? What license would they fall under?
- kmeisthax 4y agoThe license on the dataset only covers the collection of images and labels; which is not copyrightable in the US but copyrightable in the EU. The reason for this is the same reason why you can't copyright a phone book in the US but can in the EU[0]. The CC-BY license on the dataset only means you can avoid getting sued by the people who collected LAION-5B by attributing them. The copyright on the actual images and text labels is far more of a problem. Generally speaking, it is extremely infringing to collect a bunch of images or captions and redistribute them. Like, getting-to-the-heart-of-ownership, if-you-cant-sue-for-this-you-own-nothing kind of infringing. However, it's when you start talking about training ML systems on the dataset that things get interesting. In the EU, there's an explicit copyright exception for data mining that would apply to, say, training DALL-E or Stable Diffusion. But what about the US? Well, it's legal to crawl the web; and we specifically have Authors Guild v. Google where the Authors Guild lost in court trying to keep Google from scanning large numbers of books. AI researchers have sort of just taken this to mean "training AI is fair use". This is not court-tested, but it at least jives with some precedent, so I think it's OK to assume it's true. However, it means absolutely nothing for the people actually using the AI, because fair use is not transitive. If I take every YouTube video review of a movie and edit them down to just the movie clips being used, and then assemble them back together... I haven't somehow made a "fair use copy" of a movie that you can just share around. I've just made the most inefficient form of copyright infringement you can do with a computer. Likewise, if I train an AI on a movie, that can be fair use, but asking it to spit the movie back out is not. Now, keep in mind that some ML systems (such as Copilot) are very eager to reproduce their training set data. Sometimes in situations you wouldn't expect. These sorts of things are ticking time bombs for people who want to generate novel images, because the AI having trained on such a massive data set also gives you access to basically the whole data set. That's half of a US copyright infringement claim right there - the other half being substantial similarity, which basically is the "Corporate needs you to find the differences between these two pictures" meme. The only way to keep AI from infringing copyright is to make sure it never sees anything that could potentially be under copyright. [0] Strictly speaking, the EU has a separate concept of sui generis database ownership, but for this discussion we can treat it the same as copyright. If you're wondering why phone books aren't copyrightable, the term of art to search for is "sweat of the brow".
- pxoe 4y agohow would commercialization (such as, selling things made with AI, like artwork, books, etc., monetizing AI models, like selling access to them) affect 'fair use'? so far, what I'm getting is that training a model on data may be 'fair use' (though, is it creating a derivative of that data? are training results a derivative? and would there be a difference if it's made for/used for commercial purposes, or not), but the output seems to be kind of in a 'copyright limbo'.
- kmeisthax 4y agoOne of the fair-use factors is commerciality of the reuse; specifically non-commercial uses are more likely to be fair. However, this factor is treated with little weight. Practically speaking it is very difficult to imagine a reuse of a copyrighted work that does not carry some commercial benefit to someone. At the very least, not having to pay for a license is a commercial benefit of its own. Generally speaking, assume all uses are commercial and you will understand a lot of modern fair use cases. Legally speaking, "fair use" and "derivative work" are mutually exclusive. In fact, both terms were coined at the same time when SCOTUS created the entire derivative rights regime basically out of thin air in Folsom v. Marsh. They needed a legal tool to prevent people from stealing large sections of a work, but also didn't want to allow copyright owners to abolish the 1st Amendment. Hence, they set up a set of deliberately murky legal tests to determine if a use was "fair" or not. If you want a quick standard to gut-check against, the question you'd ask is: "is this use something that other people would ordinarily pay for?" If so, then it's infringing. If not, then it might be fair use. So you can see why making an image generator might be fair use, but it's output would be infringing if you could identify an original work the AI was cribbing from. It'd be difficult to even fathom how licensing on a training set would work, given that there's no clear chain of value from a particular entry in the set to a particular model weight or output. But we can clearly identify if an AI system is regurgitating training output, or has been told to copy someone else's work and change it a little - at least after-the-fact in a court of law.