18 ms·
Ask HN: DALL-E was trained on watermarked stock images?
I just got a Dall-E render with a very intact "gettyimages" watermark on it. I'm no legal expert on whether you have to own the license to something to use it as training input to your AI model, but surely you can't just... use stock photos without paying for the license? Maybe I'm just old fashioned.
Prompt: "king of belgium giving a speech to an audience, but the audience members are cucumbers"
All 4 results (all no good as far as the prompt is concerned): https://ibb.co/gz5RDkB https://ibb.co/gz5RDkB
Fullsize of the one with the watermark https://ibb.co/DzGR063 https://ibb.co/DzGR063
- cercatrova 4y agoBased on the new scraping ruling with LinkedIn [0], anything that is "open gate" (as in, accessible without logging in) can be scraped and (I assume) be used by neural networks. The onus, it appears, is to not use it to generate copyrighted works, like Iron Man from Marvel, just as one can use Photoshop as a tool but is still barred from making and selling an Iron Man digital painting. [0] https://cdn.ca9.uscourts.gov/datastore/opinions/2022/04/18/17-16783.pdf https://cdn.ca9.uscourts.gov/datastore/opinions/2022/04/18/1...
- resoluteteeth 4y ago> Based on the new scraping ruling with LinkedIn [0], anything that is "open gate" (as in, accessible without logging in) can be scraped and (I assume) be used by neural networks. The ruling you are linking to is about whether scraping violates the Computer Fraud and Abuse Act. This isn't really applicable here. First of all, that's a separate issue from copyright. Just because scraping publicly accessible data doesn't violate the CFAA doesn't mean that suddenly all images posted on the internet are public domain or that can use copyrighted images from websites for whatever you want, for example. Furthermore, how copyright applies to training neural networks on copyrighted works is an open question right now.
- cercatrova 4y agoPerhaps a neural network's outputs can be deemed as transformative use and thus fair use. I don't know though, I'm not a lawyer.
- olliej 4y agoI would assume that for cases like this is is more a matter of whether you can redistribute copyrighted work that has not had any of the usual "creative use" things applied, rather than whether the original scanning was protected.
- davikr 4y agoYeah, I've seen an image get generated with a very recognizable watermark for a certain stock image company. This happened with a totally unrelated prompt.
- dd36 4y agoDid you try reverse image searching the generated image?
- im3w1l 4y agoI remember when people used to say ianal. Innocent times when we thought there was an objective law and lawyers knew it. But that's not how these things work. The truth is that no one knows. Ultimately a bunch of people will decide how they feel about it. Well-read legal scholars trying really hard to be fair, but still just people. No one can predict with full certainty which way it will go.
- otoburb 4y ago>>No one can predict with full certainty which way it will go. Until somebody tries to float a trial balloon (case) in court.
- petesergeant 4y ago> Ultimately a bunch of people will decide how they feel about it Some would argue that technically these people _discover_[0] the law, but it amounts to the same thing [0] https://www.jstor.org/stable/3143421 https://www.jstor.org/stable/3143421
- RcouF1uZ4gsC 4y agoWhat is interesting is a human analogy. Say you were an artist who went to every art show and museum and studied all the art there. If you produced a work of art solely from memory that contained large portions of other people's copyrighted art, would that still fall under copyright/require licensing?
- whycombinetor 4y agoPrecedent in music says sometimes-yes. The "Blurred Lines" lawsuit found that Pharrell and Robin Thicke were liable in the tune of $7m for producing a work of art solely from memory that copied the "signature phrases, hooks, bass lines, keyboard chords, harmonic structures and vocal melodies" of a Marvin Gaye song. https://en.wikipedia.org/wiki/Pharrell_Williams_v._Bridgeport_Music#Holding https://en.wikipedia.org/wiki/Pharrell_Williams_v._Bridgepor... https://www.npr.org/2015/03/11/392375390/-7-million-verdict-blurs-the-lines-on-music-sampling https://www.npr.org/2015/03/11/392375390/-7-million-verdict-...
- Geonode 4y agoYes, but that's an outlier ruling that was widely criticized.
- colejohnson66 4y agoWidely criticized doesn’t change that it’s case law other judges might consider.
- alcolade 4y agoAnother human analogy could be: you take a photo from every art show and museum, and use those for reference as you paint.
- colejohnson66 4y agoThe analogy of using your cortex is more apt. My understanding of neural networks is that there are no remnants of the original inside it. The training data is used to back propagate a bunch of weights. Your brain works like those neural network neurons; they learn when to fire, but they don’t know the intricate detail like a photo. Hence why many claim eyewitness testimony is bogus.
- jcims 4y agoLegally wouldn't it just boil down to the license on the watermarked image? BTW you can add 'royalty free' to the prompt to get rid of those most of the time (lol?).
- robocat 4y ago> royalty free Wouldn’t that remove the king of Belgium? Or add a “down with the king” placard?
- purpleblue 4y agoIs there a copyright protection in terms of consuming a copyright-protected image? I thought it was only for the purpose of displaying that image. If you're reading the file and reading the data, but not displaying it, is that also protected?
- teddyh 4y agoCopyright, as the name implies, is mostly for restricting copying (as in printing copies of a book), but also restricts distribution, adaptation, display, and public performance of that work. In the case of AI, it’s the “adaptation” part which is up for debate. If a person uses an image as part of a training set for an AI image generator, and then uses said AI to generate new images, are those images “adaptations” of the images in the training set? I would suggest that the answer is yes, but current behavior by AI vendors are not in concordance with that view.
- inasmuch 4y agoWondered the same thing recently … https://news.ycombinator.com/item?id=31159231 https://news.ycombinator.com/item?id=31159231
- ShamelessC 4y ago> but surely you can't just... use stock photos without paying for the license? They aren't hosting the infringing content. Training on the data is probably covered under fair use. Generations are of _learned_ representations of the dataset, not the dataset itself. This makes it closer to outputting original works (probably owned by the person who used the model). The players involved here are known for being litigious, however. I wouldn't be surprised if OpenAI did in fact pay some hefty fee upfront to get full permission to use these images.
- kej 4y agoIf they had been paying for the images upfront, wouldn't you expect them to train the model on the non-watermarked versions?
- namrog84 4y agoThe watermarked version might be more prolific with better metatags and descriptions around them. The non watermarked versions are likely internal only and have far less diverse descriptions.
- ShamelessC 4y agoGood point! They certainly had no obligation to pay, either. Perhaps they just scraped it all.
- gricardo99 4y agoif they paid for access, or permission, why train on the watermark versions? I’m guessing they assumed fair use and there will be lawsuits.
- criddell 4y agoIs that representation of the watermark a trademark? If so, then copyright infringement might not matter, but use of the trademark may.
- BeefWellington 4y ago
- Geonode 4y agoIt doesn't matter. I could put a Getty watermark on anything. Getty would have to show that a generated image was at least in part the same as their image.
- addaon 4y agoNo. You could put the Getty watermark on anything, and that wouldn't be copyright infringement... but it would be pretty clear trademark infringement.
- angusturner 4y agoRelevant earlier discussion about this issue: https://news.ycombinator.com/item?id=32436203 https://news.ycombinator.com/item?id=32436203
- sva_ 4y agoSimilar thing with GH Copilot. I'd say it is still fair use though, even though such things should be filtered out.
- dkjaudyeqooe 4y agoA couple of people here have asserted that it's "probably" fair use, but are there any rulings on the subject?
- dlg 4y agoI am not a lawyer, but I've had to argue about copyright with several. In the United States, there are two bits of case law that are widely cited and relevant: In Kelly v. Arriba Soft Corp (9th), found that making thumbnails of images for use in a search engine was sufficiently "transformative" that it was ok. Another case, Perfect 10 (9th), found that thumbnails for image search and cached pages were also transformative. OTOH, cases like Infinity Broad. Corp. v. Kirkwood found that that retransmission of radio broadcast over telephone lines is not transformative. If I understand correctly, there are four parts to the US courts' test for transformativness within fair use (1) character of use (2) creative nature of the work (3) amount or substantiality of copying (4) market harm. I'd think that training a neural network on artwork--including copyrighted stock photos--is almost certainly transformative. However, as you show, a neural network might be overtrained on a specific image and reproduce it too perfectly--that image probably wouldn't fall under fair use. There are also questions of if they violated the CFAA or some agreement crawling the images (but Hiq v Linkedin makes it seem like it's very possible to do legally) and whether they reproduced Getty's logo in a way that violates trademarks (are they trying to use it in trade in a way there could be confusion though?)
- eslaught 4y agoSearch engines don't create market harm for a work because they don't compete with it. In fact, they do the opposite: they advertise the work, making it more accessible and increasing exposure. These AI tools on the other hand seem to do the exact opposite. They can (or could, if they got good enough) absolutely compete with a work, and therefore seem like they create substantial market harm. The character of use also seems vastly different; AI tools are creating images explicitly to be consumed, vs a search engine is basically just an index, and only shows the image in so far as it needs to make it discoverable. So three of the four tests for fair use seem clearly against AI image generation, at least to me. The only test that possibly goes in favor of AI is the amount or substantiality of copying, but AIs can easily reproduce images, or if not entire images, other substantial subsets of a composition. I just don't get how these could possibly be fair use.
- gojomo 4y ago
- chrismorgan 4y agoAll large-scale public machine learning stuff is depending on being exempt from copyright restrictions, under fair use doctrine. Look at my responses to all of the threads about Copilot + GPL for more info about that application of it: https://hn.algolia.com/?query=chrismorgan+copilot+gpl&type=comment https://hn.algolia.com/?query=chrismorgan+copilot+gpl&type=c.... When that is finally tried in court, if it fails to any meaningful extent at all (including going all the way up to Supreme Courts as it doubtless will), then Copilot is dead, DALL·E is dead, GPT-3 is dead, all of these things will be immediately discontinued in at least the affected jurisdictions, at least until such a time as they get the laws changed or judgements overturned.
- eslaught 4y agoTo me this feels like the argument that we should allow Uber and Airbnb because they're sufficiently "transformative" use cases. When clearly they are playing fast and loose by the rules and have taken advantage of being early enough to do so. As soon as the rulemakers caught up, it became obvious that they didn't have a license to operate differently from everyone else, just because they're new and popular. Personally, I agree that a strike against AI fair use would kill these current generation of tools. But I don't see why that would be the end of it. What it would do is to create a market for open source data sets with liberal licenses. We'd lose something by not being able to train models on every piece of media that has ever been on the internet anywhere, but it's not obvious to me that was ever really reasonable in the first place. If the only way to make AI that can produce good writing is to train it on every piece of writing ever produced in the history of the human race... aren't we missing something? Surely if AI has a future, it'll have to overcome this at some point.
- musicale 4y ago> As soon as the rulemakers caught up, it became obvious that they didn't have a license to operate differently from everyone else, just because they're new and popular. This is well said. One of the primary advantages of these businesses was evading the regulation and taxation that their competitors were subject to.
- webwielder2 4y agoThese are the absolute worst DALL-E images I've seen. Do people generally just share the amazing ones and most of the output is actually complete shite? Like Instagram presenting the top 1% of people's lives.
- smlacy 4y agoPrompt Engineering can help a lot but yes, you're basically right: People are generating many, many images and sharing only the best ones with the fewest artifacts. For simple prompts with little additional guidance, all the diffusion image generators I've seen/used will produce output about like what the author linked most of the time. There are always a few gems, and honing in via prompt engineering helps immensely.
- yreg 4y agoI disagree, with proper promptcrafting you can expect far far better results than the one in the op, before any cherry picking. (see my comment sibling to yours) With Dalle-2 I get a satisfying result in >50% of attempts and I'm a beginner. With Midjourney the result almost always looks great, but often misses some part of what I wanted. I'd say Stable Diffusion is similar. The results are seldom crap, but it's difficult to bend it to produce unusual situations. And in SD it's difficult to keep the entire objects in the frame, but that's a different problem.
- gojomo 4y agoOf course people are more likely to share the best iamges – or in this case, the one most illustrative of their concern (about watermarks). Also: my sense is that getting the best results often requires a lot of extra coaching with style/detail words. As we can't see the prompt here, we don't know what sort of style/details were requested. GIGO.
- tnzk 4y agoYou're right. This shows the prompt and it doesn't have such style directives https://ibb.co/gz5RDkB https://ibb.co/gz5RDkB
- gojomo 4y agoYou may want to use the native 'Share' option, especially on the one with the watermark. You'll get a public link, at `labs.openai.com` rather than some random image-sharing site, which will show the image & the prompt used to generate it (including a credit to "your-first-name × DALL·E").
- StillLrning123 4y agoKids in school are also trained on stock images https://www.reddit.com/r/KidsAreFuckingStupid/comments/8tgxsm/getty_washington/ https://www.reddit.com/r/KidsAreFuckingStupid/comments/8tgxs...
- snickerbockers 4y agoThat's technically not a stock image, it's a portait that has been public domain for a long time.
- anigbrowl 4y agoBut you've seen many PD images reshared by stock imagery companies. It raises the question of why false assertions of ownership aren't easily prosecuted, given that they constitute a kind of fraud upon the public.
- mminer237 4y agoThat's what trademark law is. But I don't think the person who painted the portrait is going to be suing Getty for impugning his good reputation. If something's public domain, anybody can use for anything they want, even if that's just rehosting it with your watermark.
- totetsu 4y agoI'm not even an American, and I've heard all about the Getty Images Address.
- jfoster 4y agoI think you're technically right, but that this will be overlooked from a legal perspective because it's less obvious that humans have been training ourselves on the prior art of others. We tend to blend in additional things besides prior art. (eg. nature, sensations, etc.)
- snickerbockers 4y agoRegardless of whether or not training an AI on stock images violates the license, there's a very real problem with that watermark being present, which is that it proves their AI is prone to copying large swaths of images from gettyimages unaltered, and that definitely is a license violation. This makes me think back to the controversy over github copilot; if these AIs are going to be trained on other peoples' IP then somebody needs to be held accountable when they commit plagiarism. Otherwise, im sure Microsoft won't mind my new "gamemaker AI" that i trained on that new halo game last year, or this "OS AI" that I trained on windows 11.
- throwaway675309 4y agoJust because it contains the text of the watermark does not mean that it's reproducing large swaths of the image - its doubtful even the most generous perceptive hash would retrieve any matching images in the Getty repository.
- snickerbockers 4y agoand yet the watermark is there. if it can't copy parts of the training images then how and why did it copy the watermark?
- kaetemi 4y agoIt's prone to copying things it has seen thousands of times, such as those watermarks. The content itself is unrelated.
- agnosis 4y agoGot the exact same girl from the picture in the ad at the bottom. Creepy! https://ibb.co/dBLNxQ6 https://ibb.co/dBLNxQ6
- deleted 4y ago[deleted]
- _trampeltier 4y agoIf you read the licence from Getty, they say, you are not allowed to use Getty pictures for ML.
- chrismorgan 4y agoWhat that license text says is irrelevant, because they’re not using it under that license, but under fair use exemptions in copyright law.
- antioppressor 4y agoFair use, until it disrupts market, sidelines creators, gobbles up the market, make it ubiquitous, lock people in, then let the regulators craft some bs law that will change nothing, compensates noone. ;)
- deleted 4y ago[deleted]
- fxtentacle 4y agoYes, Imagen and everything based on LAION 400M or 2B, too. BTW, Copilot also ignored all licenses of the source code it memorized. Datasets are the new capital. If they could, most employees would probably also object to their company using the result of their work to replace their job. But they can't. It's the same with artists here.
- Cypher 4y agoYou transformed the original enough so it's ok
- yieldcrv 4y agosome people go into business models that simply have no legal protections
- vivegi 4y agoJust wait until they build an AI watermark identifier and remover (which is a problem subset) and then use its output to train/update their model. They probably already have specialized filtering models built to filter out censorable terms. They may be imperfect, but they are there. A watermark remover might be an easy addition. When Stable Diffusion released their model playground, I used the prompt Peter at the pearly gates dressed as a security guard and got three images two of which were censored and one that was an ordinary image. So, the capability is there already. Just a matter of time before they get good at watermark removal.
- severak_cz 4y agoProbably just some stock photos with watermark sneak in. There are lots of photos with watermark circulating on web, for example in memes and unfinished webpages (when finished, these will be replaced with paid variant without watermark).
- trention 4y agoMy personal opinion is that it's unethical (and possibly illegal, in a subset of cases) to train models on data without explicit consent of the creators of that data. And that really encompasses all data - generative models were not a thing when said data was created and no matter how it was licensed before, explicit consent about using it for model training must be obtained from the creators themselves. That being said, arguments about copyright are just a fig leaf as far as I am concerned. The outcome of whether this is allowed or not will depend on the net impact of using those models on the job market and whether society will be willing to tolerate it.
- throwaway120983 4y agosometimes people will post stock images on sites with user generated content. if their training data included images scraped from those sites, then it could have gotten in that way unintentionally
- throwaway120983 4y agosome people will post images with watermarks on social media or other sites with user generated content. if their dataset included images scraped from them, then it could have gotten in that way
- xg15 4y agoReminds me of the discussion about GitHub Copilot using the entirety of GitHub as training data. I was honestly baffled how many people, even experts in the field, saw use as training data as non-infringing. With the corrolay that it's apparently perfectly legal to "copyright-wash" a work by feeding it to an AI and have that AI generate a slightly different but extremely similar work. Considering how strict and heavy-handed copyright handling has been otherwise, this has added to my belief that copyright in practice is really just enforcement of the interests of whatever industry has the most power at a given time: When entertainment and content generation was the biggest revenue generator, copyright couldn't be strict enough, now all money is on AI and suddenly loopholes the size of barn doors pop up.
- rich_sasha 4y agoWritten laws are vague, practical verdicts are based on case law, cases are won by better-funded lawyers, rich industries prevail. It's a bit of an exaggeration but maybe not too much.
- wongarsu 4y agoThese loopholes are purely theoretical until tested in court. At some point a generating AI will hurt the wrong company, and they will either make a public spectacle out of it in court, or if they see no chance of winning lobby congress to introduce laws that make the case winnable.
- xg15 4y agoYeah, things should get interesting when the first model makes use of Rings of Power or House of the Dragon footage or whatever the latest superhero movie is. I wonder if we'll see a "Hollywood vs Silicon Valley" lobbying battle. Or possibly "Amazon media division vs Amazon AI division"...
- namrog84 4y agoI think silicon valley would win. I saw some analysis a long time a go that basically indicated a couple big companies could likely buy out the entire Hollywood and music industry and fully own them and make most copyright issues go away. I dont know if its still true but its really a big difference in how much capital, revenue, and money there is. Hollywood is pennies in comparison. I think Big tech would easily win.
- JacobiX 4y agoThe first thing that I try after generating an image from DALL-E is using reverse image search. I do it on every image that I intend to use, more often than not, I find a very similar image, in this case I discard it and vary my prompts.
- supermatt 4y ago> more often than not, I find a very similar image Can you give an example? I was also doing reverse image searches, and I havent seen a single case of an image being closely related to another unless it was used as the base for inpainting.
- JacobiX 4y agoAt first, you can experiment with "extreme" input that can cause overwriting like: a photo of Marylin Monroe, etc. To search using reverse image search I use yandex, and I downsample the image.
- topicseed 4y agoWhat are the best apps and subscriptions to generate these? No private beta, just sign up, put a credit card on file, and use? (Low volume, perhaps 100 images per month, so 300-500 attempts.) Could be great for featured images for blog posts.
- surfacedetail 4y agoI'm finding it amusing that everyone immediately assumes infringement, OpenAI is a company that will not be inviting lawsuits. We can't assume any licensing behind closed doors, my guess is that OpenAI has an agreement with Getty, take a look at the licensing in this Observer piece, it's been licensed by Getty, this would indicate that Getty are happy with scraping. https://www.theguardian.com/commentisfree/2022/aug/20/ai-art-artificial-intelligence-midjourney-dall-e-replacing-artists https://www.theguardian.com/commentisfree/2022/aug/20/ai-art... Besides, this is not infringement in principle, the AI has been trained to think that high-quality news images have watermarks.
- coldtea 4y ago>but surely you can't just... use stock photos without paying for the license? You'd be surprised...
- RobertoG 4y agoI don't know about the images, but what about the watermark itself? Can I just take any photo and add a proprietary watermark?
- zlqanst 4y agoObviously you could send it to the copyright holder and find out. In the case of Copilot, Oracle certainly would sue.
- JaceLightning 4y agoEducational is a fair use category. These tools advance science. I wouldn't expect them to respect copyright.
- ratonofx 4y agoThis copyright "issues" are against the true nature of innovation. By the means of Artificial INTELIGENCE, we must to accept a mind or intelligence is free to perceive external elements and use every stimulus to execute its own creative process. The world is a perpetual iteration cycle amongst human beings. Good artists borrow, great artists steal.
- drcode 4y agoThis comment should be on the Wikipedia article for "parody indistinguishable from reality"
- Hnaomyiph 4y agoRegardless, I think I agree. I use images from all sources as inspiration for my art. I know people who have used my pieces for inspiration as well. Intelligence and creativity aren’t bound by IP law. Why should AI be bound by it?
- ratonofx 4y agoYeah! You got what I meant... We have the opportunity to make a big leap by exploring a new Era of creativity provided these technological advances. It's the same cynical people questioning if "machines going to replace people" didn't figure out that the machines need to build themselves first. When machines build themselves (with no humans since conception), let's see how the "patent ideas" world wouldn't fall apart. Until then, We can get our piece of Cake by augmenting our human creative process. Sorry but I'm not the daydreamer here. I can't live under the false premise that Intelectual Property is something that would control the input(training, learning, etc.) for all the A.I. (created or to-be-created).
- sixothree 4y agoIf we are expected to believe AI have a "creative process" then they should abide by labor and copyright laws.
- userbinator 4y agoThis interesting era of AI will surely teach us the meaning of that old phrase "great artists steal", or more subtly rephrased, "everything is a derived work".
- Asmod4n 4y agoLast time i checked you can source from whatever you want, legislation doesn't care. The last time i checked it was when colpilot got public, they could have trained it only on gpl code. The source license/copyright et all don't matter.
- registeredcorn 4y agoI don't care much for what laws say. If the only way someones service can work is by ingesting the work of someone else, without compensation, and then compete with that same person, that is wrong. If a company reverse engineers a competitors product, they still buy the product to tear it apart and figure out how it works. If a student learns from their teacher, then goes on to sell a similar kind of work as what their teacher makes, at least the student paid for the classes. This arrangement offers none of that. As long as theft is illegal, this should be. I'd call it parasitic, but it isn't; this is a parasite who's sole intent is to kill the host.
- sulam 4y agoI think it’s amusing that many commenters here are perfectly willing to defend DALL-E, but mention Copilot and the discussion looks radically different.
- tough 4y agoSo what happens if I start selling Dali like pieces?
- humaniania 4y agoSeems more likely to me that they add uploaded images into their data set and someone uploaded a watermarked image.
- BrainVirus 4y agoPeople here, as always, get hung up on legalese bullshit, but miss the overall picture. The dynamics in play is highly questionable. Countless artists and photographers put effort into creating their works. They put they work online to get some attention and recognition. A company comes along, scrapes all of it and starts selling access to the model to generate something that looks highly derivative. The original cohort of artists and photographers not only get zero money or attention from this new endeavor, they are now in competition with the resulting model. In short, someone whose work was essential to building a thing gets no benefits and possibly even gets (financially) harmed by that thing. Just because this gets verbally labeled "fair use" doesn't make it fair. Additional point: Just a few years ago a bunch of tech companies were talking about "data dignity". Somehow, magically, this (marketing) term is no longer used anywhere.
- Buttons840 4y agoI'm concerned that, and predict that, we will continue to see legal efforts from large data companies to prevent their own data from being used to train similar models. They can use our data, but we can't use theirs. Time will tell. I also fear our governments are incapable of acting on behalf of the people (non-corporations) in this matter.
- gfodor 4y agoThe law fundamentally needs to evolve. Latent embeddings of large corpuses of copyrighted works is something we are going to have to wrangle with more directly, it’s not clear to me how we even ought to want it to work in terms of the rights of the copyright holders for data it was trained on. With the release of spectral diffusion, arguably the genie is out of the bottle now, so there’s probably a ceiling on how much the law can evolve to claw back any retroactively determined rights to copyright holders.
- snarkypixel 4y agoSomewhat ironically, wasn't it openai's main mission for AI to benefit humanity?
- still_grokking 4y ago
- davidguetta 4y agoNo one fucking cares. For 1 "copyrighted" image theres a thousand free with the same quality or almost. You are wasting CO2 even discussing it