6 ms·
I am not a lawyer, but I've had to argue about copyright with several. In the United States, there are two bits of case law that are widely cited and relevant:
by dlg 4y ago
I am not a lawyer, but I've had to argue about copyright with several.
In the United States, there are two bits of case law that are widely cited and relevant: In Kelly v. Arriba Soft Corp (9th), found that making thumbnails of images for use in a search engine was sufficiently "transformative" that it was ok. Another case, Perfect 10 (9th), found that thumbnails for image search and cached pages were also transformative.
OTOH, cases like Infinity Broad. Corp. v. Kirkwood found that that retransmission of radio broadcast over telephone lines is not transformative.
If I understand correctly, there are four parts to the US courts' test for transformativness within fair use (1) character of use (2) creative nature of the work (3) amount or substantiality of copying (4) market harm.
I'd think that training a neural network on artwork--including copyrighted stock photos--is almost certainly transformative. However, as you show, a neural network might be overtrained on a specific image and reproduce it too perfectly--that image probably wouldn't fall under fair use.
There are also questions of if they violated the CFAA or some agreement crawling the images (but Hiq v Linkedin makes it seem like it's very possible to do legally) and whether they reproduced Getty's logo in a way that violates trademarks (are they trying to use it in trade in a way there could be confusion though?)
- eslaught 4y agoSearch engines don't create market harm for a work because they don't compete with it. In fact, they do the opposite: they advertise the work, making it more accessible and increasing exposure. These AI tools on the other hand seem to do the exact opposite. They can (or could, if they got good enough) absolutely compete with a work, and therefore seem like they create substantial market harm. The character of use also seems vastly different; AI tools are creating images explicitly to be consumed, vs a search engine is basically just an index, and only shows the image in so far as it needs to make it discoverable. So three of the four tests for fair use seem clearly against AI image generation, at least to me. The only test that possibly goes in favor of AI is the amount or substantiality of copying, but AIs can easily reproduce images, or if not entire images, other substantial subsets of a composition. I just don't get how these could possibly be fair use.
- gojomo 4y agoAs I see it, 3 of the 4 tests are strongly in OpenAI's favor; the 'market effect' is mixed. (1) The use is highly transformative; (2) the images used were offered to the anonymous browsing public (with watermarks); (3) the end effect of training will only retain a tiny spectral distilled essence of any individual photo, or even a giant source corpus; (4) there's a potential risk of market competition from the ultimate model output, for some uses – but that's also the most 'transformative' aspect. Getty et al could potentially just ask creators of such models not to include their images – perhaps by blocking their crawling 'User-Agent' – and it might not make any real difference in the models.
- gnopgnip 4y agoThese AI generated images are directly competing with stock images. AI tools are selling images to blogs and other customers that often would purchase stock images instead. The "character of use" is not in favor of dall-e, it is a commercial use. Copyright law does not require getty to block a user agents or ask them not to include their images. Another issue here is that removing copyright management info like a watermark is a violation of the DMCA, separate from fair use or copyright infringement. These cases have statutory damages and attorneys fees awarded.
- gojomo 4y agoWhether something is directly competing for the same business would have to be evidenced, and copyright doesn't mean protection from all possible competition - it's just one factor weighed. And fair use protects many commercial uses, too, depending on proportion/character-of-original/etc. But also, none of these images are direct, or even necessarily subtantial, "copies" of other images. The generator learned from other images – the same as any human artist might. No watermark has been removed; the bigger issue may be that the spectral watermark violates a trademark. (But, I doubt consumers are likely to be confused.)
- prox 4y agoIt’s going to be interesting what the stock companies will do. Maybe they will make their own Image Generator. Perhaps we will see a case based on the new factor that is AI. An AI is not artist; they can’t be conflated. A decent artists can churn out maybe 5-10 works if he is productive. AI can churn out by the hundreds or thousands if needed. The process also isn’t the same. Anyway it will be interesting to watch this space.
- csmpltn 4y agoPutting aside the core question of the legality of training data on licensed material - what about the false advertising/copyright aspect that comes with slapping a "GettyImages" logo on some random nonsense generated by a "neural network"?
- visarga 4y agoIt's not worth discussing about Getty so much. AI labs will collect a dataset to predict if an image is watermarked. They will crawl to index the Getty images to make sure they are not in the training set. Then retrain and in 2 months the problem is solved. They can cut out a sizeable part of the training set without problem, the model will still be good. They can also OCR the output to make sure there are no blacklisted words and use an index to skip all images that look too similar to the training data. Then the argument of copyright defenders is going to be weakened. The fact that a prompt and curation are necessary also goes against the "AI works can't be copyrighted" narrative - it's generated by a human-AI team, so human work is part of the process. The core of the issue I see is that human and AI both learn from the published media but an AI can both "see" and "draw" more than a human, so there is an important distinction there.
- csmpltn 4y agoI understand that there are (both practical and theoretical) ways to reduce the chances of an AI generating an image that has copyrighted elements in it (such as the "GettyImages" logo). I'm mostly curious about the legal aspects of having a black-box system that can - under some unknown circumstances - attach openly copyrighted or trademarked elements (such as a company logo) to a piece of work.
- srg0 4y agoIt seems it is possible to generate images which are very similar to the existing stock photos if you feed getty images' description into DALL-E. I tried it with a distinctive banana image: https://imgur.com/a/0OrIr6e https://imgur.com/a/0OrIr6e
- kriro 4y agoInteresting. Adding "stock photo" to the string generated that getty tag? That is probably the most attackable (alas easy to fix) part of the issue. It will be an interesting question how close to the original a picture has to be to be considered the same (I'm sure there's some case law) and maybe there's some new research to be done regarding how to recreate the training data images with the correct search string (I suppose one could build an ML model for that). Fun times ahead
- srg0 4y agoNo, I didn't get the tag. But I suppose that Getty metadata as well as the images were used for training.
- Dylan16807 4y ago"very similar" insofar as it's following the narrow prompt, sure. > Different runs can generate different size, orientation and placement of the bananas, as well as different shades of pink. At that point it's definitely the curation causing any possible derivation. The image generator is innocently doing what you ask in an unbiased way.
- Thorrez 4y agoThose bananas are completely different. There's no copyright infringement there. I could take a photo of a banana and photoshop it repeatedly onto a pink background. That would look just as similar, and there's no copyright problem there. You can't copyright an idea.
- srg0 4y agoImages are different, but it appears that DALL-E is inspired by the aesthetics and the layout of the copyrighted material. Another example, picking a random image from the Getty Images site. "A young parkour flips through the city,guangzhou,china, - stock photo": https://imgur.com/a/pPruwzA https://imgur.com/a/pPruwzA The images are obviously different, but it appears that DALL-E maps the getty images description to similar tone, similar perspective, similar background, and similar weather conditions. I'm sure there are thousands of possible backdrops in Guangzhou, and many ways to show a parkour flip. Even in the Google image search results there's more variance than in the output of DALL-E. So you can't copyright an idea, but you can certainly scrape a copyrighted DB with image metadata, and use it to create your own product. My point is that DALL-E itself might be a derivative work of Getty Images and thousands of other online catalogs.
- jcranmer 4y agoFrom what I understand, the actual process of fair use boils down to "the judge decides in his/her gut if the use is fair, and then writes up the analysis to justify coming to that conclusion." If you look at the recent SCOTUS opinion in Google v Oracle, you can see how two judges can look at the same facts and come to almost diametrically opposed fair use analyses. My further understanding is that generally the #1 overriding concern in fair use analysis is money, which means you're more likely to see analysis along Thomas's dissent than Breyer's opinion. In this case, let me give a fair use analysis that is going to suggest that this isn't fair. Factor 1 weighs against fair use: it's not transformative because, well, transformative is extremely narrowly interpreted against fair use. Factor 2 weighs against fair use because, well, it's factor 2 and it weighs against fair use unless the underlying copyright was paper-thin in the first place. In factor 3, it's weighing against fair use because it's not copying the minimal amount of the original work to get what it needs (it copied the watermark after all!). And factor 4 of course weighs against fair use because you're essentially creating stock images which is naturally in the exact same market that a stock image provider is in. If you wanted to write a fair use analysis that finds fair use, you'd argue instead that the work was transformative, and the amount copied also weighs in favor of fair use (thus converting factors 1 and 3 to weigh in favor of fair use). You might try to argue that it's a completely different market, but I'm incredibly skeptical that such an argument could win over both a district court and an appeals court (although Breyer's opinion in Google v Oracle did basically follow this thread of analysis, its repetition is unlikely since everyone wants to pretend that Google v Oracle has 0 impact to anything outside of software). Such an analysis is possible, but unlikely, since the unspoken factor of "could you have paid for this" tends to be the factor that wins out over everything else. Note that we are going to have a SCOTUS case in the fall that will specifically explore transformative uses in the context of fair use: Warhol v Goldsmith (https://www.scotusblog.com/case-files/cases/andy-warhol-foundation-for-the-visual-arts-inc-v-goldsmith/ https://www.scotusblog.com/case-files/cases/andy-warhol-foun...). I'm not going to hold my breath that the use will be found fair, though.
- deleted 4y ago[deleted]
- sixothree 4y ago> (2) creative nature of the work Is AI even capable of having a creative nature. All that I see is re-use of source images.