> If you require licensing fees for training data, you kill open source ML.
And likely proprietary ML as well, hopefully.
(To be clear, I think AI is an absolutely incredible innovation, capable of both good and harm; I also think it's not unreasonable to expect it to play a safer, slower strategy than the Uber "break the rules to grow fast until they catch up to you" playbook.)
I'm all for eliminating copyright. Until that happens, I'm utterly opposed to AI getting a special pass to ignore it while everyone else cannot.
Fair use was intended for things like reviews, commentary, education, remixing, non-commercial use, and many other things; that doesn't make it appropriate for "slurp in the entire Internet and make billions remixing all of it at once". The commercial value of AI should utterly break the four-factor test.
Here's the four-factor test, as applied to AI:
"What is the character of the use?" - Commercial
"What is the nature of the work to be used?" - Anything and everything
"How much of the work will you use?" - All of it
"If this kind of use were widespread, what effect would it have on the market for the original or for permissions?" - Directly competes with the original, killing or devaluing large parts of it
Literally every part of the four-factor test is maximally against this being fair use. (Open Source AI fails three of four factors, and then many users of the resulting AI fail the first factor as well.)
> If they lose, they’ll survive.
That seems like an open question. If they lose these court cases, setting a precedent, then there will be ten thousand more on the heels of those, and it seems questionable whether they'd survive those.
> To be clear, I don’t like the idea of companies profiting off of people’s work. I just like open source dying even less.
You're positioning these as opposed because you're focused on the case of Open Source AI. There are a massive number of Open Source projects whose code is being trained on, producing AIs that launder the copyrights of those projects and ignore their licenses. I don't want Open Source projects serving as the training data for AIs that ignore their license.
> "How much of the work will you use?" - All of it
That depends on the interpretation of "use", and it would be interesting to read what lawyers think. You learned the language largely from speech and copyrighted works. (All the stories, books, movies, etc. you ever read/heard) When you wrote this comment did you use all of them for that purpose? Is the case of AI different?
To be clear that's a rhetorical question - I don't expect anyone here to actually have a convincing enough argument either way.
Principles applied to human brains are not automatically applicable to AI training. To the best of my knowledge, there's no particular law that says a human brain is exempt from copyright, but it empirically is, because the alternative would be utterly unreasonable. No such exemption exists for AI training, nor should it.
Ideas/works/etc literally live rent-free in your head. That doesn't mean they should live rent-free in an AI's neural network.
Changing that should involve actually reducing or eliminating copyright, for everyone, not giving a special pass to AI.
> To the best of my knowledge, there's no particular law that says a human brain is exempt from copyright, but it empirically is, because the alternative would be utterly unreasonable.
Human brain most definitely is not exempt. If you read Lord of the Rings and then write down a new book, with the same characters and same story line - that's plain copying(lookup the etymology of the verb to copy). If you look at a painting and paint a very similar painting - that's still copying.
Human brains are the reason we have copyright. Your recital of passages from any copyrighted book would violate the copyright, if not for fair use doctrine. And it has nothing to do with whether you do it yourself, or have a TTS engine produce the sound.
The human brain is absolutely exempt, insofar as the copy stored in your brain does not make your brain subject to copyright, even if a subsequent work you produce might be. Nobody's filing copyright infringement claims over people's memories in and of themselves.
I'm saying that AI does not and should not automatically get the exception that a human brain does.
It’s not so clear cut. Many lawyers believe all that matters is whether the output of the model is infringing. As much as people love to cite ChatGPT spitting out code that violates copyright, the vast majority of the outputs do not. Those that do, are quickly clamped down on — you’ll find it hard to get Dalle to generate an image of anything Nintendo related, unless you’re using crafty language.
There’s also the moral question. Should creators have the right to prevent their bits from being copied at all? Fundamentally, people are upset that their work is being used. But "used" in this case means "copied, then transformed." There’s precedent for such copying and transformation. Fair use is only one example. You’re allowed to buy someone’s book and tear it up; that copy is yours. You can also download an image and turn it into a meme. That’s something that isn’t banned either. The question hinges on whether ML is quantitatively different, not qualitatively different. Scale matters, and it’s a difference of opinion whether the scale in this case is enough to justify banning people from training on art and source code. The courts’ opinion will have the final say.
The thing is, I basically agree with you in terms of what you want to happen. Unfortunately the most likely outcome is a world where no one except billion dollar corporations can afford to pay the fees to create useful ML models. Are you sure it’s a good outcome? The chance that OpenAI will die from lawsuits seems close to nil. Open source AI, on the other hand, will be the first on the chopping block.
>Those that do, are quickly clamped down on — you’ll find it hard to get Dalle to generate an image of anything Nintendo related, unless you’re using crafty language.
really it seems more like someone was afraid of angering Nintendo who is a corporate adversary one does not like to fight and thus it has a bunch of blocks to keep from generating anything that offends Nintendo, that does not really translate to quickly and easily stopping and blocking offending generations across every copyrighted work in the world.
> Many lawyers believe all that matters is whether the output of the model is infringing.
What I don't understand (as a European with little knowledge of court decisions on fair use): with the same reasoning you might make software piracy a case of 'fair use', no?
You take stuff someone else wrote - without their consent - and use it to create something new. The output (e.g. the artwork you create with Photoshop) is definitely not copyrighted by the manufacturer of the software.
But in the case of software piracy, it is not about the output. With software, it seems clear that the act of taking something you do not have the rights for and using it for personal (financial) gain is not covered by fair use.
Why can OpenAI steal copyrighted content to create transformative works but I cannot steal Photoshop to create transformative works?
What am I missing?
> "What is the character of the use?" - Commercial
Your first factor seems to not at all be like that which Stanford has in its guidelines[1], which they call the transformative factor:
In a 1994 case, the Supreme Court emphasized this first factor as being an important indicator of fair use. At issue is whether the material has been used to help create something new or merely copied verbatim into another work.
LLMs mostly create something new, but sometimes seems to be able to regurgitate passages verbatim, so I can see arguments for and against, but to my untrained eyes doesn't seem as clear cut.
[1]: https://fairuse.stanford.edu/overview/fair-use/four-factors/ https://fairuse.stanford.edu/overview/fair-use/four-factors/
Bear with me here. Rushed and poorly articulated post incoming...
In the broadest sense, generative AI helps achieve the same goals that copyleft licences aim for. A future where software isn't locked away in proprietary blobs and users are empowered to create, combine and modify software that they use.
Copyleft uses IP law against itself to push people to share their work. Generative AI aims to assist in writing (or generating) code and make sharing less neccesary.
I argue that if you are a strong believer in the ultimate goals of copyleft licences you should also be supporting the legality of training on open source code.
The obvious difference is that copyleft is voluntary, while having your art style stolen isn't.
If an artist approached a software developer, created a painting of them using their Mac, and said "There, I've done your job for you" you'd think they were an idiot.
This is the same from the other side. The inability to understand why that's a realistic analogy does not change the fact that it is.
> The obvious difference is that copyleft is voluntary, while having your art style stolen isn't.
What a curious type of theft where the author keeps their art and I get different art.
"> The obvious difference is that copyleft is voluntary, while having your art style stolen isn't."
This is why it is important whether you consider that infringement occurs upon ingestion or output. If it only matters for outputs, then artists have a problem, since copyright doesn't protect styles at all, see for example the entire fashion industry.
There is a saving grace though: Artists can make a case that the association of their distinctive style with their name is at least potentially a violation of trademark or trade dress, especially if that association is being used to promote the outputs to the public. This is a fairly clear case of commercial substitution in the market for creating new works in that artist's style and creating confusion concerning the origin of the resulting work.
Note that the market for creating new works in a particular artist's distinctive and named style kind of goes away upon the artist's passing. What remains is the trademark issue of whether a particular work was actually created by the artist or not, which existing trademark law is well suited to policing, as long as the trademark is defended, even past the expiration of the copyright.
Meanwhile, trademark (and copyright) also apply to the subjects of works, like Nintendo's Mario or Disney's Mickey Mouse or Marvel's Iron Man. But we don't really want models to simply be forbidden from producing them as outputs, or they become useless as tools for the purpose of parody and satire, not to mention the ability to create non-commercial fan art. The potential liability for violating these trademarks by publishing works featuring those characters rests with the users rather than the tools, though, and again existing law is fairly well suited to policing the market. Similarly, celebrities' right of publicity probably shouldn't prevent models from learning what they look like or from making images that include their likeness when prompted with their name, but users better be prepared to justify publishing those results if sued.
You can also make the (technical) argument that if you just ask for an image of Wonder Woman, and you get an image that looks like Gal Gadot as Wonder Woman, that the model is overfitting. That's also the issue with the recent spate of coverage of Midjourney producing near-verbatim screenshots from movies.
It might be appropriate though to regulate commercial generative AI services to the extent of requiring them to warn users of all the potential copyright/trademark/etc. violations, if they ask for images of Taylor Swift as Elsa, or Princess Peach, or Wonder Woman, for example.
> Fair use was intended for things like reviews, commentary, education, remixing, non-commercial use, and many other things
"many other things" has included, for example, Google Books scanning millions of in-copyright books, storing internally them in full, and making snippets available.
The basis for copyright itself is to "promote the progress of science and useful arts". For that reason a key consideration of fair use, which you've skipped entirely, is the transformative nature of the new work. As in Campbell v. Acuff-Rose Music: "The more transformative the new work, the less will be the significance of other factors", defined as "whether the new work merely 'supersede[s] the objects' of the original creation [...] or instead adds something new".
> "How much of the work will you use?" - All of it
For the substantiality factor, courts make the distinction between intermediate copying and what is ultimately made available to the public. As in Sega v. Accolade: "Accolade, a commercial competitor of Sega, engaged in wholesale copying of Sega's copyrighted code as a preliminary step in the development of a competing product" yet "where the ultimate (as opposed to direct) use is as limited as it was here, the factor is of very little weight". Or as in Authors Guild v. Google: “verbatim intermediate copying has consistently been upheld as fair use if the copy is ‘not reveal[ed] . . . to the public.’”
The factor also takes into account whether the copying was necessary for the purpose. As in Kelly v. Arriba Soft: "If the secondary user only copies as much as is necessary for his or her intended use, then this factor will not weigh against him or her"
While there are still cases of overfitting resulting in generated outputs overly similar to training data, I think it's more favorable to AI than simply "it trained on everything, so this factor is maximally against fair use".
> Directly competes with the original, killing or devaluing large parts of it
The factor is specifically the effect of the use upon the work - not the extent to which your work would be devalued even if it had not been trained on your work.
None of those arguments make sense. The output of AI absolutely does supersede the objects of the original creation. If it didn't, artists wouldn't care that they were no longer able to make a living.
Substantiality of code does not apply to substantiality of style. What's being copied is look and feel, which is very much protected by copyright.
The copying clearly is necessary for the purpose. No copying, no model. The fact that the copying is then compressed after ingestion doesn't change the fact that it's necessary for the modelling process.
Last point - see first point.
IANAL, but if I was a lawyer I'd be referring back to look and feel cases. It's the essence of an artist's look and feel that's being duplicated and used for commercial gain without a license.
That's true whether it's one artist - which it can be, with added training - or thousands.
Essentially what MJ etc do is curate a library of looks and feels, and charge money for access.
It's a little more subtle than copying fixed objects, but the principle remains the same - original work is being copied and resold.
Where this argument falls down for me is that "use" w.r.t. copyright means copying, and neither AI models nor their outputs include any material copied from the training data, in any usual sense. (Of course the inputs are copied during training, but those copies seem clearly ephemeral.)
Genuinely curious: for anyone who thinks AI obviously violates copyright, how do you resolve this? E.g. do you think the violation happens during training or inference? And is it the trained model, or the model output, that you think should be considered a derived work?
AI is a genie that you can't really stuff back into a bottle. It's out and it's global.
If the US had tighter regulations, China or someone else will take over the market. If AI is genuinely transformative for productivity, then the US would just fall behind, sooner or later.
Then let them! If another country put forward tighter regulations to help actual people over and above the state that holds them, then that is good in itself, and either way will pay for itself. Why are we worried about China or whoever taking over the market of something that we see has bad effects?
Like, we see this line everywhere now, and it simply doesnt make sense. At some point you just have to believe something, be principled. Treating the entire world as this zero sum deadlock of "progress" does nothing but prevent one from actually being critical about anything.
This would-be Oppenheimer cosplay is growing really old in these discussions.