4 ms·
Well, to be clear, I am mainly a public domain and copyright theory guy. I do keep up to date with everything here. But I also don't write laws; if I did, they
by mod50ack 3y ago
Well, to be clear, I am mainly a public domain and copyright theory guy. I do keep up to date with everything here. But I also don't write laws; if I did, they would look different.
Companies are allowed to train on copyrighted works - or, to be more precise, there is no prohibition in copyright law on them doing so. On the other hand, there are no particular protections.
The real question is to what extent an AI generator can return its copyrighted training data as a non-de minimis output. In other words, can I get one of the existing copyrighted works by putting in a prompt? This is something that AI developers are trying to avoid, but it's actuallg a pretty tricky problem.
Allowing AI output to be copyrighted is one of the worst ideas in copyright. Thankfully, the US Constitution as interpreted by the courts only allows for copyright to inhere in works of human creativity.
By the way, while I'm very excited about certain "AI" things, I have a very poor opinion of the merit of generative AI — and I mean in theory we well as how it stands today.
- xorcist 3y ago> Allowing AI output to be copyrighted is one of the worst ideas in copyright What does that mean, in more specific language? If I create a poster in Photoshop, it is under copyright? What about if I use a smart fill plugin? What about if I use a prompt plugin?
- Aloha 3y agoWhat would your ideal copyright regime look like? Mine would be different than what we have now, it'd be 20 years or artists lifetime, whichever is shorter - then 10 year long renewals are possible after that, but the cost of the renewal would ratchet up with each renewal. I've also considered using a percentage of revenue for the work - basically a tax on the revenue from that work, as a condition for the right of monopoly on it - which would also ratchet upwards with each renewal. I'd also consider a use it or lose it strategy for copyright like trademark, meaning if you are not making the work available for purchase/license within the copyright renewal period, for reasonable terms, you lose the ability to renew it. Mine is mostly designed to deal with orphaned works, ensuring they enter public domain in a predictable way, I think the biggest issue with our existing copyright system isn't enriching Disney - they're still putting those works out there, making them available - its all the works being lost to the sands of time.
- hedora 3y agoDisney regularly censors or modifies classic films in its catalog, and also simply pulls things out of production. If you want to read more, look up the Disney Vault strategy.
- Aloha 3y agoI'm well aware of Disney's business strategy. I don't care if they make their millions still. I care much more about all the works that cannot find an audience because of uncertain copyright status, and not enough commercial demand to justify figuring out who 'owns' it.
- mod50ack 3y ago> What would your ideal copyright regime look like? I would have to actually write something blog-length about this.
- bee_rider 3y ago“Find me a prompt to generate this image,” seems like an interesting problem to toss at an AI, I wonder if anyone knows of work in that direction? Hypothetically if such a system existed where you passed in a copyright image, and then got a prompt to generate it, would that be sufficient to show some kind infringement?
- dahart 3y ago> Companies are allowed to train on copyrighted works I’m curious what you mean by this. If I’m not allowed under copyright law to make a personal copy of the latest Pixar movie or watch it without permission or payment, even if I’m not sharing it with anyone else, what under the law allows a company to make a copy and train on it? I thought I understood copyright law to not only prohibit redistribution of copyrighted works without permission, but also to prohibit consumption of copyrighted works without permission? Is that accurate? In that sense, I would have thought copyright law does prohibit companies from training on copyrighted works.
- brookst 3y agoYou’re hitting on the distinction between duplication and training. I own hundreds of paperback books. Copyright law does not limit what I can learn from them. It may be that assembling a corpus for training is illegal, but if so, that would be true even if it was never used for training. The act of training an AI is orthogonal to the collection of the corpus.
- dahart 3y agoI think I’m not necessarily talking about training. It’s the duplication before training, and maybe just the fact that training is consuming the entire work, which I think under copyright law requires permission (which typically means payment). > It may be that assembling a corpus for training is illegal, but if so, that would be true even if it was never used for training. Yeah, exactly! You’re right that copyright law doesn’t limit what you can learn at all, and doesn’t copyright ideas. But it does, I think, limit whether you’re allowed to read the copyrighted work in it’s entirety the first place, if you haven’t paid for it or legally borrowed a copy or whatever. Gaining access to the material is covered under the law, right? This does mean, I suspect, that assembling a corpus of copyrighted training material is not allowed under copyright law, unless it was all paid for or licensed with permission. If the AI companies have paid for all the material they used to train, then my question might be moot, I’m assuming they didn’t pay for it. This is murky when there’s a lot of copyrighted material that’s available online, maybe with the intent that it would be consumed in small parts and not copied wholesale by machines for the sole purpose of making software that can replicate the content and style of what it learned.
- BeefDinnerPurge 3y agoAI models have an amazing ability to approximately memorize any training data. It's just that that memorization is useless unless it memorizes something real as opposed to random (randomized ImageNet labels as a simple example). So as much as I want there to be a fair use case here, the artists have a real point. If someone can break the memorization without losing significant validation/test set performance, that might go a long way. But even then, artists don't want their style copied either, and that's problematic to me in that if a human does it, that's OK, but if an AI does it, it's not? Yes I get the ease of asking an AI to do it vs a 10K+ hours artist, but, well, more or less the same to me on a geological time scale. In the next year, I'm hoping to Patreon/Kickstart project that offers two major funding tiers. Hitting the lowest tier means it will use AI to create assets, and hitting the higher tier will use humans instead. My response to this brouhaha is to throw the controversy right back at the people creating it in the first place and ask if they're willing to walk their fancy talk on this subject with their wallets.