8 ms·
It's not that the watermark is on them per se, but that the model tried to emulate an image it had seen before which had a watermark on it. Imagine showing a ch
by cheald 4y ago
It's not that the watermark is on them per se, but that the model tried to emulate an image it had seen before which had a watermark on it. Imagine showing a child a bunch of pictures with Getty watermarks on them, then they draw their own, with their own emulation of the watermark. They don't know it's a watermark, they don't know what a watermark is, they just see this shape on a lot of pictures and put it on their own. That's essentially what's going on.
The model is only around 4GB, and it was trained on ~5B images. At 24 bit depth, that'd be 786k raw data per image, which would be 3.5 petabytes of uncompressed information in the full training set. Either the authors have invented the world's greatest compression algorithm, or the original image data isn't actually in the model.
So, I think the argument is: if you look at someone else's (copyrighted) work, and produce your own work incorporating style, composition, etc elements which you learned from their work, are you engaged in copyright infringement? IANAL but I think the answer is "no" - you would have to try to reproduce the actual work to be engaging in copyright infringement, and not only do these models not do that, it would be extremely hard to get them to do so without feeding them the actual copyrighted work as an input to the inference procedure.
- wnevets 4y ago> It's not that the watermark is on them per se, but that the model tried to emulate an image it had seen before which had a watermark on it. Imagine showing a child a bunch of pictures with Getty watermarks on them, then they draw their own, with their own emulation of the watermark. That's essentially what's going on. The blurred watermark is what makes its obvious they used Getty's (copyrighted?) images to train the model.
- sdenton4 4y agoA world where we can't use copyrighted material to update neutral network weights is a world where we can buy books but not read them...
- wnevets 4y agoIt is allowed to sale a book that is collection of pages from copyrighted books? Paragraph 1 is from a Stephen King novel, Paragraph 2 is from A Storm of Swords and so on? I am not a copyright attorney but that sounds like violation to me.
- PeterisP 4y agoIf you have legally obtained copies of the relevant novels, according to the first sale doctrine you should be allowed to cut them up, staple first chapter of one novel to the second chapter of another and third chapter of the next, and then sell the result. But the authors have the exclusive right to making more copies of their work, so if you'd want to make a thousand of these frankensteinbooks, you would need to get a thousand copies of the original books.
- snovv_crash 4y agoYou can train it but not for commercial purposes. Nobody cares what you do at home, but if you want to use someone else's work to make money they will come knocking for their cut.
- sdenton4 4y agoDoes O'Reilly ask for a percentage of a software engineer's income after they've read their book of perl recipes?
- snovv_crash 4y agoO'Reilly sells their books to software engineers with the intent for them to use the information to further their knowledge and apply it in a commercial setting. The images in Getty are provided with the intent to be used only as a catalogue for purchasing corresponding images without watermarks. The difference in intent is very clear, and a judge would make a distinction between these.
- cheald 4y agoI understand that, but why is using copyrighted images to train a model be any more illegal than studying copyrighted paintings in art school? Copyright doesn't prevent consumption or interpretation, simply reproduction.
- joe_the_user 4y agoIf a student studied an older master in school and produced a painting inspired by that old master that included a copy of the signature of the old master, this would be more indication of intent to fraud than if they didn't include the signature. Copyright can be fairly flexible in interpreting what constitutes a derivative work. The Getty water is evidence that an image belongs to Getty. If someone produces an image with the watermark and gets sued, they could say "your honor, I know it looks like I copied that image but let's consider the details of how my hypercomplex whatsit work..." and then judge, say a nontechnical person, looks at the defendant and say "no, just no, the court isn't going to look at those details, how could court do that?". Or maybe the court would consider it only if you paid 1000 neutral lawyer-programmers to come up with a judgement, at a cost of millions or billions per case.
- filoleg 4y agoWhat if it was not the copied signature of the old master, but a new one with a similar style and placed in a similar spot on the painting, but with the name of the student instead and looking blurry/a bit different? Because that's what's happening here, and that doesn't sound quite like fraud. Another scenario, what if i create a painting of a river by hand in acrylic and also draw a getty-watermark-looking thing on top using acrylic? As for why, i would put it there as an integral part of the piece, to allude to the fact of how corporations got their hands over even the purest things that have nothing to do with them, with the fake watermark in acrylic symbolizing it. You can make up any other reason, this is just the one i thought of as i was writing this. It wont look exactly like the real getty watermark, it will be acrylic and drawn by hand, so pretty uneven with colors being off and way less detailed. Doesn't feel like fraud to me.
- 4y ago
- PeterisP 4y agoNoone is contesting the fact that images where copyright is owned by Getty were used the model. The contested issue is whether training a model requires permission from the copyright holder, because for most ways of using a copyrighted work - all uses except those where copyright law explicitly asserts that copyright holders have exclusive rights - no permission is needed.
- sdenton4 4y agoIt's not even necessarily trying to emulate any particular image it's seen before; it may just decide 'this is the kind of image that often has a watermark, so here goes.'
- notahacker 4y agoI think you'd struggle to argue that the Getty watermark was a general style and composition principle and not a distinct motif unique to Getty (and in music copyright cases, the defence of plagiarising motifs inadvertently frequently fails).
- cheald 4y agoFrom the model's perspective, it's not a distinct motif, that's the thing (and, it struggles quite a lot to reproduce the actual mark). The model doesn't have any concept of what a "watermark" is. As far as it's concerned, it's just a compositional element that happens to be in some images. Most "watermarks" Stable Diffusion produces are jumbles of colorized pixels which we can recognize as being evocative of a watermark, but which isn't the actual mark. A quick demo: I fed in the prompts "a getty watermark", "an image with a getty watermark", and "getty", and it spat out these: https://imgur.com/a/mKeFECG https://imgur.com/a/mKeFECG - not a watermark to be seen (though lots of water). I was then able to generate an obviously-not-a-stock photo containing something approximating a Getty watermark, with the prompt "++++(stock photo) of a sea monster, art": https://imgur.com/a/mNC6XtQ https://imgur.com/a/mNC6XtQ - the heavily forced attention on the "stock photo" forces the model to say "okay, fine, what's something that means stock photo? I'll add this splorch of white that's kinda like what I've seen in a lot of things tagged as stock photos" and it incorporates that into the image as a to satisfy the prompt. We can easily recognize that as attempting to mimic the Getty watermark, but it's not clearly recognizable as the mark itself, nor is the image likely to resemble much of anything in Getty's library.
- notahacker 4y ago> From the model's perspective, it's not a distinct motif, that's the thing (and, it struggles quite a lot to reproduce the actual mark). The model doesn't have any concept of what a "watermark" is. The court delivers the judgement, not the model. If courts can find against musicians whilst accepting they 'unconsciously' plagiarised key elements of a song in their own completely different song played by different musicians based on maybe hearing it in the background somewhere, they can certainly find against the creators of a model which has a sufficiently strong and obvious dependency on Getty IP they imported to output reasonably close approximations of Getty watermarks.
- mr_toad 4y ago> Either the authors have invented the world's greatest compression algorithm, or the original image data isn't actually in the model. AI and finding the best compression algorithm for an input are essentially the same problem.