9 ms·
I've seen a lot of confidence on HN and other tech communities that a court would never rule that training an AI on copyrighted images is infringement, but I'm
by ar-nelson 4y ago
I've seen a lot of confidence on HN and other tech communities that a court would never rule that training an AI on copyrighted images is infringement, but I'm not so sure. To be clear, I hope that training AI on copyrighted images remains legal, because it would cripple the field of AI text and image generation if it wasn't!
But think about these similar hypotheticals:
1. I take a copyrighted Getty stock image (that I don't own, maybe even watermarked), blur it with a strong Gaussian blur filter until it's unrecognizable, and use it as the background of an otherwise original digital painting.
2. I take a small GPL project on GitHub, manually translate it from C to Python (so that the resulting code does not contain a single line of code identical to the original), then redistribute the translated project under a GPL-incompatible license without acknowledging the original.
Are these infringements?
In both of these cases, a copyrighted original work is transformed and incorporated into a different work in such a way that the original could not be reconstructed. But, intuitively, both cases feel like infringement. I don't know how a court would rule, but there's at least some chance these would be infringements, and they're conceptually not too different from distilling an image into an AI model and generating something new based on it.
- stjohnswarts 4y agoHonestly, as much as I am rooting for AI "art" (still not sure about that term here) I can see how Getty would easily have a claim in court if the AI was indeed trained on some of their images AND they can prove it somehow. If that's not derivative then I don't know what is. Maybe a special niche could be carved out for people who are only researching and experimenting and not really "selling" or profiting from the resulting images. It would seem if they're right that maybe they could bury a watermark in their images that identifies it as Getty (or just whomever) and that it's copyrighted by them and they don't give permission to use it for training AI. Maybe I just don't know enough about how the algorithms work though shrug
- nneonneo 4y agoIt’s not about reconstruction, it’s about the notion of a “derivative work”. Translating a work would absolutely be derivative (consider the case of translating a literary work between languages: this is a classic example of a derivative work). Blurring a work but incorporating it would nonetheless still be derivative, I think. The challenge with these models is that they’ve clearly been trained on (exposed to) copyrighted material, and can also demonstrably reproduce elements of copyrighted works on demand. If they were humans, a court could deem the outputs copyright infringement, perhaps invoking the subconscious copying doctrine (https://www.americanbar.org/groups/intellectual_property_law/publications/landslide/2016-17/march-april/ghosts-hit-machine-musical-creation-doctrine-subconscious-copying/ https://www.americanbar.org/groups/intellectual_property_law...). Similarly, if a person uses a model to generate an infringing work, I suspect that person could be held liable for copyright infringement. Intention to infringe is not necessary in order to prove copyright infringement. The harder question is whether the models themselves constitute copyright infringement. Maybe there’s a Google Books-esque defense here? Hard to tell if it would work.
- dTal 4y ago> The challenge with these models is that they’ve clearly been trained on (exposed to) copyrighted material, and can also demonstrably reproduce elements of copyrighted works on demand. If they were humans, a court could deem the outputs copyright infringement, perhaps invoking the subconscious copying doctrine (https://www.americanbar.org/groups/intellectual_property_law https://www.americanbar.org/groups/intellectual_property_law...). Every single human has been exposed to copyrighted material, and probably can reproduce fragments of copyrighted material on demand. Nobody ever writes a book or paints a picture without reading a lot of books and looking at a lot of paintings first. For a "subconscious copying" suit to apply, you need to demonstrate "probative similarity" - that is, similarity to copyrighted material that is unlikely to be coincidental. In other words - it's not clear to me that the situation with AI is any different than with a human, or that it presents new legal challenges. If it looks new, it is new.
- judge2020 4y agoNote that a derivative work doesn't instantly make it 'fair use' for the purposes of copyright. You typically still need 'adaptation' permission from the copyright holder to made a derivative work of it, so you can't make 'Breaking Bad: The Musical' by recreating major scenes in a play format, at least not without substantially changing it[0]. For the purpose of fair use, copyright.gov has an informative section titled "About Fair Use" which details what sort of modifications and usage of a copyrighted work would be legal without any permission from the copyright holder https://www.copyright.gov/fair-use/#:~:text=a)(3).-,About%20Fair%20Use,-Fair%20use%20is https://www.copyright.gov/fair-use/#:~:text=a)(3).-,About%20... 0: https://en.wikipedia.org/wiki/Say_My_Name!_(Musical) https://en.wikipedia.org/wiki/Say_My_Name!_(Musical)
- indymike 4y ago> I've seen a lot of confidence on HN and other tech communities that a court would never rule that training an AI on copyrighted images is infringement, but I'm not so sure. To be clear, I hope that training AI on copyrighted images remains legal, because it would cripple the field of AI text and image generation if it wasn't! Regardless of the copyright of the training data which really is unresolved, the copyright-ability of output of AI is questionable at best. There's no way to monetize the generated images for a stock art that isn't at risk of a court ruling pulling the rug from under it.
- ummonk 4y agoAI-produced art is still human-made, as a person does the job of engineering a prompt and selecting from the generated images. The copyrightability of such work is unlikely to ever seriously be in question.
- deleted 4y ago[deleted]
- ska 4y agoThat's not so obviously clear cut. Can the model produce identical output from the same simple prompt?
- MintsJohn 4y agoBut that's exactly what happens, AI isn't randomness, it's a set of predefined calculations. The randomness is in the seed/starting point. For e.g Stable Diffusion it is given/user input, resulting in perfect reproducibility.
- ska 4y agoI should have been clearer - that was rhetorical to point out that if you and I use the same prompt and pick the same resultant image, it's harder to claim either of us have copyright. This is a gray area. The tools themselves could introduce some stochastic aspect so that outputs are never identical, also. None of this leads to an obviously clear cut legal position wrt copyright.
- joe_the_user 4y agoI've seen a lot of confidence on HN and other tech communities that a court would never rule that training an AI on copyrighted images is infringement, but I'm not so sure. Indeed and your examples are intended to point to gray areas. But a much more problematic (for the user) example is: some Dall-E-like program spits out seemingly original images but 0.1% are visibly near duplicates of copyrighted images and these form the basis of a lawsuit that costs someone a lot of money. Copyright in general tends to use the concept provenance - knowing the sequence of authors and processes that went into the creation of the object [1] and naturally AI makes this impossible. The AI training sequence either muddies the waters hopeless or creates a situation where the trainer is liable to everyone who created the data. And I don't think the question will answered just once. The thing to consider is that anyone can just create an "AI" that just spits out stock images (which is obviously a copyright violation) and so the court would have to look at detailed involved and neither the court nor the AI creator would want that at all. ianal... [1] https://serc.carleton.edu/serc/cms/prov_reuse.html https://serc.carleton.edu/serc/cms/prov_reuse.html
- aaroninsf 4y agoUnfortunately for the Getty et al, "feels like" is worth exactly $0. Any court that understands the technology at even a lay level, has no path to find infringement applying existing precedent. NB "style" is not protected.
- eropple 4y agoI do understand the technology at at least a lay level (unless your Scotsman is true by your conclusion), and it seems to me that the idea that it's "style" and not "permuting the input data" is one that seems to be a postulate, not a fact.
- Beldin 4y ago> To be clear, I hope that training AI on copyrighted images remains legal, because it would cripple the field of AI text and image generation if it wasn't! To be clear, there's no law banning training an AI. There are laws for what you can do with other people's stuff. In short, maybe the AI field would indeed be crippled if they no longer freely take input from others without asking permission and/or offering compensation. And maybe that's far, far from a bad thing.
- ar-nelson 4y agoThat's true, but AI models trained on copyrighted images already exist and can't just be removed from the internet, and their output will often be indistinguishable from that of "clean" models. What I fear is a kind of legal hazard that would make even the possibility that AI had been used anywhere in a work radioactive. Imagine another hypothetical: I create a derivative work by running img2img on another artist's painting without their permission. Whether the AI model in question contains copyrighted content or not, this is probably infringement. Now suppose that, instead, I create an original work, without using img2img on someone else's art. But, as part of my process, I use AI inpainting, with a clean AI model, so that the work has telltale signs of AI generation in it. And then suppose an artist I've never heard of notices that my painting is superficially similar to theirs--not enough to be infringement on its own, even with a subconscious infringement argument. But they sue me, claiming that my image was an img2img AI-generated derivative of theirs, and the AI artifacts in the image are proof. With enough scaremongering about AI infringement, it might be possible for a plaintiff to win a frivolous lawsuit like this. After all, courts are unlikely to understand the technology well enough to make fine distinctions, and there's no way for me to prove the provenance of my image! If it becomes common knowledge that AI models can easily launder copyrighted images, and assumed that this is the primary reason people use AI, then the existence of any AI artifacts in a work could become grounds for a copyright lawsuit.
- kmeisthax 4y agoBoth hypotheticals are likely infringement. The first example may be considered de minimus, but the courts hate using those words, so they might just argue that you didn't blur it enough to be unrecognizable or that it could be unblurred. However, the thing that makes AI training different is that: 1. In the US, it was ruled that scraping an entire corpus of books for the purpose of providing a search index of them is fair use (see Authors Guild v. Google). The logic in that suit would be quite similar to a defense of ML training. 2. In the EU, ML training on copyrighted material is explicitly legal as per the latest EU copyright directive. Note that neither of these apply to the use of works generated by an AI. If I get GitHub Copilot to regurgitate GPL code, I haven't magically laundered copyrighted source code. I've just copied the GPL code - I had access to it through the AI and the thing I put out is substantially similar to the original. This is likely the reason why Getty Images is worried about AI-generated art, because we don't have adequate controls against training data regurgitation and people might be using it as a way to (insufficiently) launder copyright.
- LtWorf 4y agoSearching books is different than generating books and selling them.
- afro88 4y agoIt might technically be infringement, but the proof is in the pudding. It may be very hard to prove a specific image (or set of millions of images) were used in training.
- kranke155 4y agoIt’s completely ridiculous to believe that copyright claims are unenforceable because you ran it through an ML transformation engine. I hope Getty sues and wins. Train your datasets on your own data! This is mass IP theft.
- cheald 4y agoNonsense. Observations about certain characteristics of a copyrighted work are not covered under that work's copyright. If I take a copyrighted book and produce a table of word frequencies in that book, no serious person would claim that the author's copyright domain extends to my table.
- Xelynega 4y agoIs the copyright still not applicable if your encoding table has rules to reconstruct text from? No serious person would argue that copyright doesn't extend to the encoded version of the book and prevent me from profiting off it. I believe the same applies to AI generation of text/images. Just because you're encoding the data in statistical models doesn't mean that it's not encoded.
- kranke155 4y agoEveryone in HN keeps pretending like ML transformation = human inspiration. This is really funny - we don’t have AGI but we have an AGI-like capability to avoid copyright. Seems to be the only place where human rights and AI rights are matched is where it most benefits AI research. How interesting. A for profit computer program =! A human being.
- dTal 4y agoCan you articulate a meaningful (and more importantly legally provable) difference? Both human brains and these sorts of AI programs are intractable black boxes. Maybe you think computer programs lack some sort of divine spark, but even if we accept that it seems to me that it's not a given that humans apply their divine spark every time they create something either.
- not2b 4y agoFor case 2, translating a novel from, say, English to Japanese, still requires permission of the holder of the copyright to the English version, even though the resulting novel "does not contain a single line ... identical to the original".
- kromem 4y agoThere's a literal arms race behind the scenes in AI right now. I think it's very unlikely that corporate IP claims that could substantially hold back progress in domestic AI development will end up being successful.
- TaylorAlexander 4y ago> To be clear, I hope that training AI on copyrighted images remains legal, because it would cripple the field of AI text and image generation if it wasn't! Temporarily, yes. However it wouldn’t be the worst thing if they were forced to work on sample-efficiency. And I maintain that if they want a large dataset of images they can collaborate with Twitter and Instagram to add an image license option to the image uploader, and provide a license option that explicitly allows this kind of use. AI needs sample efficiency research and if they actually had to get consent for their images there are a lot of artists who wouldn’t feel like their work was being ripped off. It’s probably better if this kind of use is broadly permitted, but I don’t think it’s the disaster some think it would be if they were forced to get consent. There’s already millions of openly licensed images in several collections online (Creative Commons, Wikimedia Commons) and if they forced people to provide licenses on sites like twitter we could actually have open datasets for this instead of these kind of grey areas of privately scraped datasets.
- deleted 4y ago[deleted]
- dcow 4y agoI don't know if there's a clear answer to (1), but with respect to (2) I believe there is precedent that the copyright owner would have a strong case if they had reason to believe the python library author had seen their work and it played a roll in the transcription to python. You can't look at a GPL work and literally transcribe it to license launder. Companies have tried. You can implement a similar idea in a clean room, though.
- iroh2727 4y ago> because it would cripple the field of AI text and image generation I don't disagree with this statement. But arguably, it's becoming clear that these fields exist, on an economic level, as a means for already powerful corporations and technocrats to gain ownership over and repurpose the labor of previous generations for profitable automation (not simply to create C-3PO or something). Not unlike all those other non-digital areas of the economy (e.g. railroads were built on the blood, sweat, and tears of previous generations and they are now owned by a small few, likely unrelated to the descendants of the laborers who built them). For some ML-related commentary on this subject, see e.g. https://nathanieltravis.com/2022/08/01/ai-research-the-corporate-narrative-and-the-economic-reality/ https://nathanieltravis.com/2022/08/01/ai-research-the-corpo...