4 ms·
Whether it's fair use or not (under US law) is still very unclear to me. Consider the "fair use" criteria laid out in Folsom v. Marsh: 1) the purpose and chara
by cle 3y ago
Whether it's fair use or not (under US law) is still very unclear to me. Consider the "fair use" criteria laid out in Folsom v. Marsh:
1) the purpose and character of the use, including whether such use is of a commercial nature or is for nonprofit educational purposes
This is quite muddied. OpenAI is in some commercial and non-profit superposition, and many of its users are commercial and using the technology for commercial applications. But a huge swath of its users are using it for nonprofit educational purposes too. I use it primarily for learning, along with most people I know who use it. IMO there's no clear characterization here, given the information we have. Maybe a court could compel more information to clarify this.
2) the nature of the copyrighted work
I don't know enough about it to have an opinion.
3) the amount and substantiality of the portion used in relation to the copyrighted work as a whole; and
This is also unclear to me, because the models are effectively word-lossy compression, so are unlikely to reproduce substantial portions of works that impact the copyright holder. But they could in some cases.
4) the effect of the use upon the potential market for or value of the copyrighted work
Also unclear. I can easily imagine scenarios where the use of an LLM has a negative market impact on a copyright holder, but I can also imagine scenarios where it has a positive impact (ex "Give me some book recommendations"). What's the net impact? No idea.
- bayindirh 3y agoWell, I was discussing Large Language Models (LLMs) as a technology, and esp, via this video [0] with a friend, just now. He made a striking remark. LLMs compress the information they ingest in a variably-lossy way, and when you query them, they rebuild a representation from whole or part of this compressed data in a statistical manner. As a result they're a lossy storage medium. Not unlike some music formats which store audio data in a lossy way. You know, this data can be restored with mathematical and statistical wizardry. You don't get everything you put in, but 97%-99% out of it, and companies and RIAA went bonkers for years, because even if it's not an exact reproduction, it was a reproduction enough, and this is a copyright violation. If an LLM can reproduce what I have written, or coded with 97%-99% accuracy without any license information, and I licensed this thing with less than permissive licenses, and sue the maker of that LLM, what will happen? - Will it be a copyright infringement? - Will it be fair use? It's the first if you look at it fairly, but it'll be probably be the latter, because money, fame and other corporate points will be at stake, otherwise. [0]: https://www.youtube.com/watch?v=zjkBMFhNj_g&start=323 https://www.youtube.com/watch?v=zjkBMFhNj_g&start=323 Edit: I forgot to add the video. :)
- shkkmo 3y agoYou seem to be conflating two things. Is the model itself a reproduction or is it capable or making reproductions? I think there is a strong argument for bot treating a ML model that is a mathematical amalgamation of a huge variety of material and is thus transformed as a new work. Now, if you use this system to make a reproduction of one of the works it was trained on, that doesn't "wash" the IP, that reproduction would still face all the same tests for infringement as any other work.
- danaris 3y agoIs an MP3 file itself a reproduction of the audio? Or is it merely capable of producing the reproduction? Now, it's trivially true that an MP3 file is not, itself, physically audio. It is physically bits in a digital storage medium. Without the right software and the right commands, there is no way to reproduce the recorded audio from the MP3 file. Now, that software is pretty ubiquitous today, so we have come to think of MP3 files as being the audio...but in a technical sense they do actually share a lot with the models behind LLMs and other generative ML projects. In another 10 years, the software to reproduce elements from the training set of an LLM may be as ubiquitous as MP3 players are today. (Well, that's actually pretty unlikely, given how many devices can play MP3s, but such software could well be baked into the OSes of major desktop and mobile operating systems by that time, anyway.) I think it is an oversimplification to call an LLM a "lossy compression" of the training data, but being an oversimplification doesn't mean there's not some very real truth to it—possibly enough that the law would find it to be a reasonable analogy.
- hn_acker 3y ago> Is an MP3 file itself a reproduction of the audio? Or is it merely capable of producing the reproduction? Regardless, when the MP3 file is played using a program following the MP3 standard, the result will be a sound identical or substantially similar to the sound that was encoded into the MP3 file. The purpose of the MP3 standard is to encode audio into a file (which will be called an MP3 file) which when played using a program following the MP3 standard produces a sound similar to the original audio. An AI model is meant to be created by aggregating multiple works, but the purpose of the model is not necessarily to produce something substantially similar to any work in the set. In my framing from the previous paragraph, the purpose of an AI model that produces images is to produce an image that can be opened using JPG/PNG/WEBP/AVIF image viewer, where the new image isn't necessarily supposed to be but can be similar to an existing work. If you were to train an AI model on literally only one work, then you would get something analogous in purpose to an MP3 file, since why would you use the model to produce something nonsimilar to the single work in the training set? Pasting what I wrote in a different comment: My argument is that similarity of new works (including AI model outputs) to actual works (not to non-existing works, existing styles, or theoretical aggregations of multiple works) is a prerequisite to infringement. From an article about the substantial similarity test in the US [1]: > To win a claim of copyright infringement in civil or criminal court, a plaintiff must show he or she owns a valid copyright, the defendant actually copied the work, and the level of copying amounts to misappropriation.[1][3] In order to get an infringing output, the user usually has to include a reference to an existing author or an existing work in the prompt. Sometimes that's not the case (which I've worried about with respect to Copilot), but in order to damn the model as a whole, you would have to establish that in over some percentage of cases the model produces infringing outputs for prompts which don't reference a particular author (whether individual or collective), a particular work, or a style strongly associated with a single or a few authors. [1] https://en.wikipedia.org/wiki/Substantial_similarity https://en.wikipedia.org/wiki/Substantial_similarity
- ToucanLoucan 3y ago> OpenAI is in some commercial and non-profit superposition This is the big issue I have with it that no one has yet had a satisfactory (to me anyway) answer to. The internet was scraped for all manner of writing, images, etc. to train these models for research. And then once the research was done (enough), they took the model and began selling it for profit. The question of is using publicly available content as training data fair use is an interesting question, but OpenAI has gone to market with a product on offer not merely without answering it, but seemingly quite deliberately avoiding answering it and AI hype people seem completely fine with it, despite that being a make-or-break question for the concept. And it's less that I think OpenAI owes royalties to every single person they collected training data from, and more than I'm extremely uncomfortable with and skeptical of the notion that, from my perspective, yet another major player in the tech space is abusing the public square to make a buck, which is unfortunately far from a new story. Maybe I'm just getting old, but the rallying cry of disruption, where these startup companies come in half-cocked insisting an old industry is in dire need of innovation but plot twist, the innovation is just app-powered slavery again, they get seventy billion dollars to build an app and a website, crush an existing industry underneath the weight of VC money, then either a) price the service so it's no goddamn cheaper than the thing it replaced but now the people actually doing the work are somehow making even less money or b) go out of business and leave a husk of an industry left that barely functions and nobody wants to start it back up again is just... hollow. Technology is great, but these massive corporations with the backing of unholy amounts of money just manifesting their destiny all over our society and economy and making us live with the consequences isn't.
- pseudalopex 3y agoThe nature of each use matters. Not the nature of the user. Not the nature of the user's customers. The effect of each use matters. Not the net effect of all possible uses.
- wredue 3y ago>This is also unclear to me, because the models are effectively word-lossy compression, so are unlikely to reproduce substantial portions of works that impact the copyright holder. But they could in some cases. If you tried to submit schoolwork with some of the minimal amount of changes that even the best models spit out, you’d be expelled dude. Contrary to the implication of your statement, language models don’t actually understand what the words they spit out actually mean. They can regurgitate definitions.