4 ms·
The dataset this is built from (https://www.gwern.net/Danbooru2018 https://www.gwern.net/Danbooru2018) is, simply put, copyright-infringing on a gross scale. Th
by tofof 7y ago
The dataset this is built from (https://www.gwern.net/Danbooru2018 https://www.gwern.net/Danbooru2018) is, simply put, copyright-infringing on a gross scale. The vast majority of the images uploaded to 'boorus' completely lack a compatible license or the artist's express consent. The redistribution via torrent of 2.5 TB of some 3 million images only compounds this problem. None of this is ameliorated by the $20 'generosity' of the dataset creator.
As a result, every single artist whose work was included in that dataset has a clear, meaningful claim that each and every 'waifu' sold ($20, if customized, or $5 if random) by Sizigi Studios is an infringing derivative work. Coupled with at least one of the project authors' ready admissions -- in this very comment section -- of scraping image sites himself, I would say that this team is playing with fire. Even in the case that an algorithm's output is somehow found to be 'creative' rather than mechanistic, AND this specific application is found to be in all cases substantially transformative, there's STILL the original massive 2.5 TB of copyright infringement up front to deal with.
All an enterprising lawyer would need to begin is to search the BigQuery metadata for the 'artist' and 'copyright' tags on these images. Note of course that the 'copyright' tag is widely misused on boorus and similar image repositories to refer to the inspiring franchise; 'trademark' would be much more accurate descriptor.
EDIT:
I do not mean to suggest that litigation from the use of the dataset in this ML (as opposed to the original, clearly infringing, download & redistribution) would in any way be an easy, one-sided case --- only that this scenario would represent nearly the worst possible test case imaginable for determining the future legality of ML, short of directly antagonizing the RIAA or MPAA.
- bdon 7y agoInteresting. How would you respond to an argument that the output of a ML model is transformative / fair use? Would your argument apply to this specific application, or to ML in general?
- tofof 7y agoObviously the question of whether an ML model is adequately transformative is an unanswered one, at least as far as I'm aware, in any jurisdiction. However, in this specific case, I would expect most courts to place heavy weight on the clear, initial massive-scale infringment from the dataset alone to conclude that no good-faith effort (or, apparently, any effort at all) was made to avoid trampling these artists' rights. Such 'dirty hands' would potentially discredit any attempt to claim original creative expression, rather than commercialism, motivated the creation of this ML model. This represent very nearly the worst possible case to serve as a potential test case for the legality of ML techniques. Other datasets, like the Open Images Dataset (https://arxiv.org/abs/1811.00982 https://arxiv.org/abs/1811.00982) explicitly recognize and address this concern in their curation of included images.
- wbl 7y agoIt's a standard part of artist training to copy paintings by hand, infringing copyright. Never to the best of my knowledge has this been used to argue that a picture with no visible elements of another infringes. If these infringe, where is the similarity?
- tofof 7y agoTrue, and an interesting philosophical question. However, it is not a standard part of artist training to obtain and redistribute, without license, the (in this case millions) of paintings they studied. > Never to the best of my knowledge has this been used to argue that a picture with no visible elements of another infringes. One only needs to look as far back as 2013, in Williams v. Bridgeport Music, to find such a thing not merely argued, but successfully litigated. In this case, the estate of Marvin Gaye alleged that Robin Thicke's "Blurred Lines" copied the 'feel' and 'sound' of "Got to Give It Up" despite containing no samples or even an identical chord progression. Perhaps more surprising to you will be the fact that the court found in favor of Gaye's estate, i.e. that "Blurred Lines" was infringing! A not-insubstantial factor in reaching this decision was, as I alluded to above, the attitude of the defendant regarding the infringement. Thicke testified "No" when asked if he considered himself an honest person, and admitted that "Got to Give It Up" was a direct inspiration for the song. It would be difficult to argue that the data used to create your ML was anything BUT its explicit inspiration, and as I've mentioned in other posts, this is compounded by the fact that the acquiring the initial dataset is itself a separate and very clear-cut case of copyright infringement. In the interest of good discourse, do note that many legal scholars and industry experts were, admittedly, shocked by the decision and decry it as fundamentally mistaken. Nevertheless it is now certainly precedential caselaw.
- Grue3 7y ago> If these infringe, where is the similarity? It would be quite interesting to compare the generated images with the dataset using a tineye-style image matcher. I wouldn't be surprised if large segments of the generated pictures are outright identical to some image from the dataset.
- ipsum2 7y agoThe ImageNet dataset, which is probably the most commonly used dataset (besides MNIST), is also "copyright-infringement on a gross scale". http://image-net.org/about-overview http://image-net.org/about-overview
- meruru 7y agoFWIW Danbooru complies with artists requests to remove all of their work.
- tofof 7y agoSuch compliance may protect Danbooru itself from some specific causes of action, particularly if Danbooru is able to convince the court that they were truly innocuous middlemen and that the site itself was not designed from the ground up to facilitate such infringement. Such a task would be an uphill battle in most jurisdictions subject to the Hague Convention. However, all of that is irrelevant when the concerns are the distribution of the site's contents by torrent, or the use of those copyrighted images in making derivative works. In both cases the party in question is now first-party to the concern, either the distribution or the derivation.
- lawrenceyan 7y agoThe model they use is simply a large matrix of numbers. What copyright infringement is there?
- jchw 7y agoTo be fair... that is also an apt description of a 2D image.
- lawrenceyan 7y agoSure, but in the case of the copyrighted work, we're talking about a specific combination and sequence of numbers which represent that 2-D image. The numbers here representing the weights that make up the model have nothing to do with those.
- jchw 7y agoYes, I agree that there is an argument to be made here, just not on the basis that the model is 'just a matrix of numbers.' Of course, when we're talking about this many weights, one must wonder exactly how much of the input images is actually pretty much directly encoded in the resulting network.
- throwaway99111 7y agoCan one make a one-to-one correspondence of such numbers to a copyrighted work of art?
- lawrenceyan 7y agoThe way these types of networks function results in the generation of wholly unique, never before seen images. I guarantee you will never find a single copyrighted work of art that is generated by a model like this.
- throwaway99111 7y agoIt seems then that the closest analog is the network is a derived work then.
- jchw 7y agoI think this is a phenomenally interesting point of view. It is true that 'boorus' do not bother with artist consent, and I think by and large the sentiment from Japanese artists has been fairly negative, if a bit muted (perhaps this is a bit of a cultural thing?) Despite that, the dynamic between boorus and content creators has always been complicated. Boorus have some positives even for artists. For one thing, they organize images pretty extensively; it is not abnormal to see a booru post with 50 to 100 tags, and it's probably fairly uncommon to see less than 10 tags - so they're great for searching for images. Because of that, they are an excellent place to search for inspiration or references, and I definitely know folks that do this. They also act as content aggregates, which does help people discover content and artists, especially combined with tags. Boorus do tend to comply with "do not post" requests, though that obviously does not mean artists are then implicitly consenting to their work being posted or used. But of course, the boorus themselves are not charities. Boorus tend to run ads or even accept direct payment from users in exchange for features. Personally, I do find this unethical, and I highly recommend you do not browse a booru without an adblocker, not even because of ethical concerns but due to the fact that many of them have pretty nasty advertising (Also, be aware that many boorus, like Gelbooru for example, are not particularly shameful about explicit content, and neither are their ads.) I was involved in a small scale booru years ago in my youth, for a particular interest. It's funny, I was so wrapped up in the utility of organizing and categorizing images that I had not even considered the issue of artist consent or copyright (I was not making any money, in my case, at the very least.) Ultimately, an artist complained that their work was repeatedly posted without proper links, something we did our best to discourage, and I shut the whole thing down in short order, for better or worse. Like many things on the internet, boorus would not be nearly as useful without their flagrant disregard for copyright. It's not just booru owners in this debate, though - some online artists take a view that copying and sharing online is both inevitable and the point, while others, probably moreso these days due to an increase in bad actors, take a hardline stance against copyright infringement. The law clearly and plainly sides with the latter, and I think I do too. Before sites like Pixiv existed, there were very few resources for finding Japanese illustrations in one place, much of it scattered around in fc2 blogs and Geocities sites (Yes, even until pretty recently! Geocities Japan outlived many other Geocities regions.) Nowadays, there's not nearly as bad of a discovery problem with Japanese art, and the boorus feel more parasitic than they once did. (Aside: I think many software engineers are used to being very technical about copyright, even when they do not understand the details correctly. Most artists I know do not dabble much into the details of copyright or licensing, and may even have some difficulties grasping the implications of say, Creative Commons terms. It's important to recognize that not everyone has the same point of view and things that seem obvious to us may not even make sense to others.) Of course, though, that a booru is basically massive copyright infringement is one thing. Now we're talking about neural networks built from them. And damn, that is complicated, and probably breaking new ground. I'm sure people with strong opinions will claim that it is very black and white, but I can't see this as anything less than a dilemma. I agree the model clearly constitutes copyright infringement if redistributed, but the results... that seems deeply complicated. On one hand, for all of their impressive strides as of late, current neural network algorithms are not really that impressive in their creativity. It's clear they are very heavily influenced by the data sets. That said, though... at the end of the day, our brains are also 'neural nets.' If we design a sufficiently advanced neural net and feed it a handful of pictures from Danbooru, and it is able to spit out similar but clearly distinct art, how is that any more copyright infringement than a human artist that draws inspiration from the same images? I think people predicted that the copyright situation would get complicated when DeepDream showed up, but it's nothing compared to the calamity that could occur as a result of this kind of data set.
- stale2002 7y ago> All an enterprising lawyer would need to begin is to search the BigQuery metadata for the 'artist' and 'copyright' tags on these images. Maybe, but they probably won't right? Like, if you had to bet money on if anyone is going to sue these guys over the next couple years, which side would you bet on? I'd bet on "No". The reality is that these people probably made some X0,000 dollars on this, and nobody is going to bother to pay lawyers a bunch of money to go after something, that only kinda sorta looks like infringement to machine learning experts who are squinting really hard, all in order to claim their 1/3millionth percentage claim on it. I don't see anybody wasting their time and money to go after a machine learning waifu vending machine that is hosted at a couple anime conventions.
- avian 7y agoI think it's not as clear as you claim and more like a big gray area, like a lot of fan art out there. Some images in the dataset themselves might be infringing on some other publisher's rights. On the other hand some artists knowingly submit their original art to the image boards. Where exactly is the point between a derivative work and an original work that was inspired by something? A lot of fan art I see clearly depicts a character from some franchise in a style that is close to the original, but say, in a new pose or setting. Is that copyright infringement? Trademark infringement? What if it's an original character in the exact style from the franchise? Some artists sell this kind of art as their own at conventions and will aggressively try to remove reposts on the web. On the other extreme, I've seen other artists tag any fan art remotely connected with some franchise with "copyright by <franchise owner>" and denounce any rights to their work.
- userbinator 7y agoEven in the case that an algorithm's output is somehow found to be 'creative' rather than mechanistic, AND this specific application is found to be in all cases substantially transformative, there's STILL the original massive 2.5 TB of copyright infringement up front to deal with. You're almost making the argument that everyone who has ever read a pirated book or other work (I know plenty of people who started their whole career that way...) and then uses that knowledge is guilty, which is just a perfect example of how ridiculously insane copyright law is. "Everything is a derivative work." "Stand on the shoulders of giants."
- aperrien 7y agoThe claim also appears to extend to everyone who has in their life ever read a library book.
- astrange 7y agoThe anime art industry in Japan is based on a copyright-deténte system where new artists publish copyright violating fanart at real life events like Comic Market, which is a good enough business that it supports them until they become the next generation of professional artists. Of course, it helps that some series like Touhou have explicit copyright licenses allowing this. But basically everything's OK unless you try to make porn for a Nintendo game.