5 ms·
So some images were overrepresented in the dataset, and subsequently, the network overfitted. Known problem, known solution.
by usrbinbash 4y ago
So some images were overrepresented in the dataset, and subsequently, the network overfitted. Known problem, known solution.
- cool_dude85 4y agoThe thread specifically says that deduplication, based on their own attempts with smaller models, helps but is not sufficient to prevent this.
- haswell 4y agoThe problem and solution are independent of the implications of this finding and how those implications are likely to influence the legal landscape. To me, the implication is that these models cannot be seamlessly exchanged for a human brain when considering their impact and compatibility with current laws.
- gpderetta 4y agoHuman brains can produce near duplicates of things they have seen in the past. Of course, like an human, an NN can be used to generate copyright infringing images, but it doesn't follow that any generated image is infringing.
- haswell 4y ago> Human brains can produce near duplicates of things they have seen in the past. Most humans cannot, and this is important, because the fact that most humans cannot arguably played a role in the formulation of all current rules. So can cameras. And there are laws that restrict where cameras can be pointed in places that do not restrict human sight, because the two kinds of "seeing" are distinctly different. Layering AI into the mix doesn't suddenly remove or mitigate the technical realities of the software. > but it doesn't follow that any generated image is infringing. I agree, and this is not my claim.
- pixl97 4y agoLaws very rarely restrict where cameras can be pointed, and most those laws focus on commercial redistribution. They almost never stop individuals from doing so. Furthermore the laws tend to be a complete wreck of logical paradoxes that fall apart in situations like this, hence forcing more case law to be generated.
- haswell 4y agoI think the reason those restrictions exist is more important here than the rarity of the restriction. I argue this because I think the magnitude of the implications of unrestricted ingestion of public and private data is similar to the magnitude of the implications of unrestricted camera use without limit, i.e. as good a candidate for a restriction as those camera use cases currently restricted by law. The reason you cannot record what's going on in a bathroom has as much to do with the implications of the recording as it does with the implications of being observed in the first place. I don't disagree about the ensuing mess of laws, but I don't think we have a reason to believe Stable Diffusion will be spared from it.
- unusualmonkey 4y agoWhy? Are you saying humans are incapable of accidentally reproducing previously seen work?
- 988747 4y agoMost humans are not capable of exactly reproducing previously seen work even on purpose. Every human artist has unique style that shows even when they try to imitate. That's why perfect forgery is a form of art in itself :)
- wahnfrieden 4y agoCorrect
- haswell 4y agoMany of the core arguments that claim Stable Diffusion training is exercising fair use do so by claiming that this computer program is similar enough to a human brain to qualify it for the protections described by copyright law, and they base that claim of similarity on the idea that the software doesn't copy, it just learns, and that all output is fundamentally new. I'm arguing that a finding like this harms such core arguments, and highlights just one of many ways these models are entirely unlike humans. > Are you saying humans are incapable of accidentally reproducing previously seen work? To conclude that would be a form of propositional and possibly equivocation fallacy, IMO. While I acknowledge that it's possible someone might "accidentally remember" someone else's work and then create a piece that is very similar to it, this hardly seems likely as a general case, and such a possibility was baked in to the current rules as written. To take this further and claim that a human could do so with precision and in a reproducible manner seems questionable. The Stable Diffusion equivalent of this kind of "remembering" and resulting duplication is again different in context and contents than a human exposed to the same images. The fact that it can be replicated systematically is the most distinctly non-human part, and is a strong hint that we're comparing very different things.
- usrbinbash 4y ago> I'm arguing that a finding like this harms such core arguments Why? If a really good artist studies a single piece long enough, he could be able reproduce it to a degree where it takes expert analysis to determine which is the original. It's not as if there have never been forgeries of expensive artworks. The difference between a human studying a certain piece intensely, and a model overfitting to it, is that to the model, it happens by accident. Overfitting to the training set is not a desired outcome, it's something ML techniques are trying to actively avoid.