6 ms·
Someone on the machine learning reddit asked me this: > Question: How does copyright work for GAN output? If I input 300,000 copyright protected photos of cele
by tasdfqwer0897 7y ago
Someone on the machine learning reddit asked me this:
> Question: How does copyright work for GAN output? If I input 300,000 copyright protected photos of celebrities and generate images of new celebrities that do not exist, are the generated images public domain or would there be copyright issues?
AFAIK, this is not a settled issue, but I'd be really interested hear what an actual lawyer thinks about this?
- mikehollinger 7y agoI’m curious as well. I imagine the definition of “derivative work” might be an interesting sub-problem to resolve while trying to answer the original question.
- tasdfqwer0897 7y agoSo if you wave your hands enough, it seems like maybe there's an argument to be made that the weights of a trained GAN somehow correspond to a 'compilation' of the training data as it's defined in this doc: https://www.copyright.gov/circs/circ14.pdf https://www.copyright.gov/circs/circ14.pdf
- genai 7y agoWith a compilation, you are able to find the original sources, but with a GAN is it even possible to find the original sources based on an output result alone?
- genai 7y agoHow would you even prove “derivative work”? My understanding is that in proving derivative work you would need to show the original, but in a GAN output which could be made up of 20MM nodes you would not even be able to confirm which 200K image(s) where used in the production of the output result
- tasdfqwer0897 7y agoThis actually might have interesting connections to ideas from differential privacy. Maybe the work is derivative of a particular training image if we can easily predict the presence or absence of that training image given only the trained model?
- genai 7y agoIf you load 50k celebrity images into a tensor of size (500000, 28, 28, 3) and then generate a resulting tensor that results in a (28, 28, 3) tensor where each of the pixel locations is merged to form an average face image (similar to https://www.dailymail.co.uk/femail/article-1355521/Average-female-face-The-Face-Tomorrow-Mike-Mike-project.html https://www.dailymail.co.uk/femail/article-1355521/Average-f... ), then although the trained model contains all the images; I would have thought the output average face image is an entirely new creative work?
- genai 7y agoGood question, IANAL but how is a human artist drawing a face (based on their learned reality of what a face looks like) any different than a GAN drawing a face (based on the GAN learned reality)?
- dragonwriter 7y ago> IANAL but how is a human artist drawing a face (based on their learned reality of what a face looks like) any different than a GAN drawing a face (based on the GAN learned reality)? A human artist is a legal person, a GAN is not. That's a fairly substantial legal difference.
- genai 7y agoBut the output result is the same (i.e. human can produce a drawing of a face, GAN can produce a new drawing of a face). If another animal such as chimp draws a human face what do you think would be the outcome? Would a chimp drawing a face be any different to a human drawing a face?
- dragonwriter 7y ago> But the output result is the same Law is only extremely rarely concerned with only output results and not process, status of actors, etc. > Would a chimp drawing a face be any different to a human drawing a face? Yes, legally, the outcome wouldd be different (whether it was original, a direct copy, or something made by copying elements but with some new content) because chimps are neither legal actors that can create a copyright through authorship nor ones who can violate copyright.
- genai 7y agoInteresting, also what do you mean by "status of actors"?
- mikehollinger 7y agoThere’s a really interesting adjacent problem - the so-called “monkey selfie” which was recently settled by stating that works created by a non human we’re not copyrightable. See https://en.m.wikipedia.org/wiki/Monkey_selfie_copyright_dispute https://en.m.wikipedia.org/wiki/Monkey_selfie_copyright_disp...
- 6gvONxR4sf7o 7y agoThe main question is that of whether the produced image is a derivative work. The interesting requirement is that the derivative work has originality. I can't tell whether there's a legal definition of originality, or if it's an "I know it when I see it" kind of thing. What's so interesting with ML is that there's a dense spectrum between memorization and originality. I hope in the future people start checking how much their models are memorizing. My favorite case is Google's facts, like when you google "golden retriever weight." From other people going out, measuring it, writing it up, and publishing that info to the web, Google can extract the info and never direct traffic to those sites. I still don't know whether I think it's okay.
- genai 7y agoFacts are not copyrightable Link: http://www.dmlp.org/legal-guide/works-not-covered-copyright#facts http://www.dmlp.org/legal-guide/works-not-covered-copyright#...
- 6gvONxR4sf7o 7y agoLegal question aside, it's still can't decide if I feel like it's dickish. Even that has a spectrum, from "What is harry potter book is first?" to "What is the first line of harry potter?" to "What is the text of the first harry potter book?"
- notahacker 7y agoThat does depend on jurisdiction and the definition of fact (for an extreme version the other way round: for a long time a UK copyright troll called Football DataCo claimed copyright of lists of UK football scores and demanded license fees - even where publishers actually obtained the scores from sending their own journalists to grounds. This was eventually overruled by the European Court of Justice, so it might be back again in future...) A lot of Google's answers to questions veer towards being written opinions and original definitions anyway
- gwern 7y agoGAN images are, I think, as far as humans go, pretty creative. It's easy to look at them knowing what they are and dismiss them (especially if you focus on the worst samples), but if you didn't know... One of the things that has most amused me about creating https://www.thiswaifudoesnotexist.net https://www.thiswaifudoesnotexist.net is watching the reactions to images being posted elsewhere, especially without (initial) attribution: not a few people like the faces enough to request info about the character or artist, or say that it looks a bit machine-learning-like but is obviously too good to actually be ML-generated, or compliment OP on their illustration! People have begun casually using them as avatars, and there's one account on Pixiv for just uploading TWDNE or other faces generated by my StyleGAN models.
- dlg 7y agoI am not a lawyer, but like many Internet entrepreneurs, I’ve had to learn a bit about copyright. There are two cases in the 9th Circuit that speak to this: Perfect 10 vs Amazon and Kelly vs Arriba Soft. As I understand it, the cases established a multi-part test for whether a particular use of images is infringing. One of the key parts is whether the use is “transformational”—I would argue that most GAN output would fall under this fair use exception for transformational use and thus would be ok (at least here in Calif). In the aforementioned cases, thumbnailing a la Google images was sufficiently transformational so synthesizing whole new images certainly should be. The tests do, however, depend on the use so I wouldn’t say that re-synthesizing similar-but-different images with a NN is always and automatically fair use.
- avinium 7y agoReally interesting question, and one that won't be settled until it's explicitly legislated (or taken to court). I have a background as a lawyer (albeit not in IP) and currently working with ML, and I can see a case being made either way.
- Isinlor 7y agoTake a look here "Why Is AI Art Copyright So Complicated?": https://www.artnome.com/news/2019/3/27/why-is-ai-art-copyright-so-complicated https://www.artnome.com/news/2019/3/27/why-is-ai-art-copyrig... > Claims that AI is creating art on its own and that machines are somehow entitled to copyright for this art are simply naive or overblown, and they cloud real concerns about authorship disputes between humans. The introduction of machine learning as an art tool is ironically increasing human involvement, not decreasing it. Specifically, the number of people who can potentially be credited as coauthors of an artwork has skyrocketed. This is because machine learning tools are typically built on a stack of software solutions, each layer having been designed by individual persons or groups of people, all of whom are potential candidates for authorial credit.