6 ms·
I get why artists are trying to stop them, but this battle has already been lost. These tools have been out too long, open source models proliferate freely, and
by srackey 3y ago
I get why artists are trying to stop them, but this battle has already been lost. These tools have been out too long, open source models proliferate freely, and jurisdictions that don’t care about IP law will continue publishing these models. By the time it works it’s way through the courts the situation will be even worse.
- tavavex 3y agoYou can publish these models in all jurisdictions, including ones that do care about IP law. There's no rulings saying that models trained on datasets of images are direct derivatives of the original images (in a way that's copyright-violating), let alone whatever is produced using these models. Now, this isn't settled law by any means, but it feels like people give too much credence to all the theft accusations by assuming that it's some slam dunk case that can make all generative AI suddenly vanish.
- kmeisthax 3y agoI used to agree with this, but there's been research coming out of Google that has altered my opinion. Specifically, Google's gotten rather good at making AI spit out unaltered training set data[0]. This is only possible if the AI is remembering large portions of the original trained-on works, which would make the weights infringing. [0] In the most egregious case, they found that just asking ChatGPT to repeat a word over and over again will make it spit out training set data verbatim. OpenAI's response was to stop any conversation with over a page of repeated words, so you can't replicate this.
- chii 3y ago> This is only possible if the AI is remembering large portions of the original trained-on works which is fine in my books. You could equivalently produce the same works from searching thru the digits of pi. The only enforcement that's needed is on the end user of the model - if they choose to produce the training set, they are violating copyright. The creator of the model _does not_ violate copyright merely by creating and distributing the weights of a neural network, as long as the space of potential output is vastly larger than the training set.
- johnnyanmac 3y ago>You could equivalently produce the same works from searching thru the digits of pi. in the same way you can brute force an MD5 Hash if you had a few centuries, I guess. I don't think "monkey's making Shakespeare" is a good metric for when to determine if a piece of art is unique enough to be scraped. Tracing is looked down upon in the art community and this feels way too close to tracing if they just have these images in a database ready to reference. You can't store copyrighted movies nor music on such private databases (one of the few times I will ever utter the words "thank you DMCA"), I don't see why art pieces would be exempt.
- bjt 3y agoThat's a really weird argument. Copyright is a legal system created by the Constitution and statutes and administrative rules. It cares about whether you are "copying", and it cares about whether you're creating things that compete with the works of the original authors. It doesn't care about potential output spaces. In this context, I don't see a principled difference between the model weights and really good compression. If I send you a gzipped copy of the latest bestseller book it's still copyright infringement. And it would still be infringement if I shipped it inside a software program that can _also_ reshuffle the words in a bajillion different ways, if there's a "copy" of the original work in there.
- scheeseman486 3y agoYou can eke out large chunks of works from Google Books with the right queries too.
- kmeisthax 3y agoYes, but Google Books is a search engine. It doesn't write books, it just tells you where a particular phrase might occur in those books. There's explicit caselaw allowing you to do this, extending back before Google was even a thing. For related reasons, Google Books also does not let you read the whole book - just the page the search match came from. OpenAI and other large language model developers are claiming they have a machine that can write books, but they also fed it shittons of books, and they can't account for where all that text went. At best they can say "well, it doesn't produce exact, verbatim copies of the training set all the time".
- tavavex 3y agoI'm kind of confused - you've claimed that Google's AI was broken, but cite an anecdote over ChatGPT? Regardless, even if these cases did happen often enough, it's erroneous to assume that this is something universal (i.e. can be applied to all generative AI models) or intentional. Said models are vastly smaller than the sizes of their training datasets, so it's more or less impossible for all the data to be stored verbatim. Some aspects may appear to look like memorization if the piece of data reappears many times in the dataset - in those cases, the algorithm has a really strong incentive to recite that data. This reduces the effectiveness and is considered an artifact that needs to be corrected, not some underlying idea of AI. The way these algorithms work is public knowledge, there's not really any black boxes in the hands of OpenAI that would be relevant here. Considering that, I'm really doubtful that these claims can be easily supported.
- jncfhnb 3y agoGoogles research on chat gpt
- numpad0 3y agoThere has been successful prompt engineering attempts to force GPT-based models to regurgitate original dataset texts texts texts texts texts texts as for example Wikipedia is a free-content online encyclopedia, written and maintained by a community of volunteers, collectively known as Wikipedians, through open colla--
- junofan 3y agoI think you’re a bit confused—these models are trained by copying works, and the memorization is the incontrovertible proof. Imagine if you download a font and use it to make a logo without paying to license the font.
- kmeisthax 3y agoGoogle does research on both their own and other people's models all the time. Here's the blog post and paper if you want to know more: https://not-just-memorization.github.io/extracting-training-data-from-chatgpt.html https://not-just-memorization.github.io/extracting-training-... You appear to be refuting a slightly different point, though. When OpenAI was making Dall-E 2, they found that duplicates in the training set would incentivize memorization of specific images, as you said. This is more like finding a secret cheat code that would make the model regurgitate everything it had remembered, regardless of how "incentivized" it is to do so or even if it had been aligned to not do that. My personal argument is this: the primary metric that the training process attempts to minimize is perplexity. This is how good the model is at guessing the training set data. Base models are specifically being designed to compress huge amounts of text and we just so happen to accidentally get a decent word calculator out of it. The alignment fine-tuning that happens later adjusts the model to prefer answering questions, but the underlying memorized data is still there. >The way these algorithms work is public knowledge, there's not really any black boxes in the hands of OpenAI that would be relevant here. Nope. Modern GPT is entirely a black box. OpenAI stopped publishing model weights the moment Elon Musk stopped writing the checks. Hell, they don't even publish the model architecture anymore. How GPT-4 works, even on a basic "this is how many transformer layers and attention heads we're using" basis, is a trade secret. Even if you have model weights, the actual meaning of the learned model parameters has never been known; there is active research on figuring them out. One particular problem is polysemanticity. If you look inside a particular hidden layer, you won't see a single "dog" or "cat" neuron in its hidden layers. You'll have a 512-dimension concept bouillabaisse with "dog", "cat", "bird", "guinea pig", "kangaroo", and so on all floating around whatever shape made sense at training time (even if it implies absurdities like "desk is the opposite of loin cloth"). To untangle this, you have to train another AI to pick out monosemantic clusters of neurons that can then be inspected, and that requires extreme amounts of GPU resources.
- skydhash 3y agoLockpicking tools are widely available, and designs of locks can be found on the web if someone search hard enough. But entering another's home is illegal. It's why businesses pay for licenses even if cracked softwares exists. I hope that artists win these and make using or creating an illegal model an high-risk activity, not worth it for any commercial activity.
- newZWhoDis 3y ago[flagged]
- coldbrewed 3y agoA society without art is a society without joy or self awareness. A society with only AI created art is a facsimile of joy or self awareness. If we're trying to remove the parts of society that challenge and inspire us then we'll be left with cultural mush that doesn't challenge, critique, explore, or even play.
- dkasper 3y agoI don’t think the current tools are doing much to the creation of “art”, it’s mostly replacing run of the mill graphic design work like logos and ads and cover art, or maybe some higher end stuff like automating workflows.
- dralley 3y agoYou'd best hope the day AI comes for your job is a long way away. Karma is, after all, a <redacted>
- dorkwood 3y agoIf anything, the AI image-generation movement proves just how highly society values the work of the artist, since we now have tech companies dumping hundreds of millions of dollars into training models to approximate their work.
- wds 3y ago
- vunderba 3y agoYup, and even if the law decides against using copyrighted images as training data, companies such as Adobe are already ahead of the curve with generative systems like firefly which has been trained exclusively on licensed artwork.
- doctorpangloss 3y agoIt is a misconception that Adobe's models have not been trained on copyrighted work. Nobody should be repeating their marketing claims. Adobe has not shown how they train the text encoders in Firefly, or what images were used for the text-based conditioning (i.e. "text to image") part of their image generation model. They are almost certainly using CLIP or T5, which are trained on LAION2b, an image dataset with the very problems they are trying to address, C4 (a text dataset similarly encumbered) and similar. I welcome anyone who works at Adobe to simply answer this question of how they trained the text encoders for text conditioning and put it to rest. There is absolutely nothing sensitive about the issue, unless it exposes them in a lie. So no chance. I think it's a big fat lie. They'd have to have made some other scientific breakthrough, which they didn't. Using information from https://openai.com/research/clip https://openai.com/research/clip and https://github.com/mlfoundations/open_clip https://github.com/mlfoundations/open_clip, it's possible to investigate the likelihood that using just their stock image dataset, can they make a working text encoder? It's certainly not impossible, but it's impracticable. On 248m images (roughly the size of Adobe Stock), CLIP gets 37% on ImageNet, and on the 2000m from LAION, it performs 71-80%. And even with 2000m images, CLIP is substantially worse performing than the approach that Imagen uses for "text comprehension," which relies on essentially many billions more images and text tokens.
- deleted 3y ago[deleted]
- joshspankit 3y agoAs well: the desire to have art seen and recognized (or even monetized) puts it in the public eye in a way that can be used (legally or not) to train AIs. Going the opposite way means committing to not showing your work to humans either.