5 ms·
So-called "fair use" for editorial and review purposes seems like the social good you're talking about? I'm not sure OpenAI offering a service for profit falls
by colonelspace 2y ago
So-called "fair use" for editorial and review purposes seems like the social good you're talking about?
I'm not sure OpenAI offering a service for profit falls into that category.
- gruez 2y agoThe supreme court ruled in Authors Guild, Inc. v. Google, Inc. that google scanning entire books and showing snippets was fair use. I don't see why AI training wouldn't also qualify.
- colonelspace 2y agoI don't see why it would qualify as fair use. OpenAI are building a product to offer to the public for profit. If I employ 10,000 humans to read books and provide summaries or texts "inspired by" those books, I need to pay for the copies of the books those humans read.
- gruez 2y ago>OpenAI are building a product to offer to the public for profit. So was google books. >If I employ 10,000 humans to read books and provide summaries or texts "inspired by" those books, I need to pay for the copies of the books those humans read. IANAL, but that would be perfectly legal. Summaries aren't copyrightable, and if you can acquire the book free but legally (eg. library, borrowing from a friend, buying it from a store and then returning it), there's nothing the publisher can do.
- prewett 2y agoThe summary is also copyrighted by the author, but it is a different copyright than the book. Probably also a bit thinner copyright than the book. Mere collections (notably, an alphabetized ordering of name and corresponding telephone number) is effectively not copyrighted, since it requires no creativity. A summary, though, requires quite a bit of creativity: what to emphasize, how to compress the ideas and arguments into a concise statement, etc.
- krapp 2y agoPossibly because companies using AI don't train them on copyrighted material simply for research or the public good, but to turn a profit on content generated by those models, and often explicitly in the style of the creators of that content. It's difficult to argue that, for instance, training a model on all of Frank Miller's work then prompting it to generate comic art in Frank Miller's style then selling that is fair use.
- gruez 2y agoAnd google books was "simply for research or the public good"?
- krapp 2y agoAs far as I'm aware, Google didn't create new versions of those books based on that content. Even if you want to argue that fair use didn't apply to Google, it clearly applies far less to what AI is used for.
- gruez 2y ago>As far as I'm aware, Google didn't create new versions of those books based on that content. It's _copy_right. If reproducing verbatim snippets was "transformative" enough to fall under fair use, I don't see why producing whole new books would not count as "transformative" enough. Copyright is a regime to grant monopoly over a specific work, it's not a regime to prevent competition from others in general.
- stillold 2y agoThere is an interesting argument here. If Google was selling brand new books created only by taking snippets from other books, that would also fall under fair use?
- gruez 2y agoThey would, but openai isn't explicitly doing that either. You can probably smuggle a full copy of a book via snippets, but that doesn't put google on the hook for copyright infringement. Likewise if openai makes you jump through hoops to produce works verbatim they should be in the clear.
- bmitc 2y agoOn a side note, does anyone use Google books? I haven't found useful information with it in a long time.
- longdustytrail 2y agoEtymologists use it a lot. I went down an etymology rabbit hole a while back looking for the origin of a phrase and google books was immensely helpful
- BeefWellington 2y agoThat's not quite the same. The argument in that case was that Google did that not in order to copy and distribute books, but rather to provide search and indexing of certain words within the text. It was important that it wasn't for the same purpose as the original works. There were also limitations on the amount that would be shown. In the case of LLM training, it's for the same purpose as the source material -- to generate code, or writing, or photographs, etc. Not only that, but in several instances it's been shown to reproduce source material, which is either derivative work or straight copying, depending. They're different situations.
- gruez 2y ago>In the case of LLM training, it's for the same purpose as the source material -- to generate code, or writing, or photographs, etc. Not only that, but in several instances it's been shown to reproduce source material, which is either derivative work or straight copying, depending. If it's used in a reference/"inspiration" capacity (as opposed to verbatim copying), I doubt the rightsholder have anything to stand on here. Sure, their works might have been used to make other competing works, but all art is derivative, and I don't see why it would be legal for a human artist to "train" on past works of art but not AI. Alleging that AI models can reproduce some works verbatim is probably the stronger argument, but AFAIK you have to coax them pretty hard to do so, and therefore AI companies might be able to argue they're tools like photocopiers or such. Likewise, you can probably extract an entire book off google books by bruteforcing common ngrams to get the entire book, but google wouldn't be held liable for that.
- BeefWellington 2y ago> If it's used in a reference/"inspiration" capacity (as opposed to verbatim copying), I doubt the rightsholder have anything to stand on here. Sure, their works might have been used to make other competing works, but all art is derivative, and I don't see why it would be legal for a human artist to "train" on past works of art but not AI. The "all art is derivative" line is essentially something people try to convince others of to justify breaking copyright law. It's not grounded in reality or law. It devalues creative work by implying the machine, with no lived experiences, is doing the same thing. And it's also completely wrong about what the specific term derivative work actually means in the context of copyright. Derivative works deal in specific. If your LLM reproduces a substantial portion of the story beats from Jurassic Park, you can bet it'd wind up in court. If it reuses identifiable characters, that is usually gonna be derivative unless it can otherwise qualify under an exemption. "But fanfic, fanart, etc." Is a common counterpoint but misses the commerce aspect of it. Here Open AI and similar are offering paid services based upon harvesting all of this information. When they produce for you the response to the prompt, they are, effectively, distributing that to you for money. That's the point at which it becomes a problem. As an aside, it's an act of drinking the LLM Kool-Aid to believe it can be "inspired". > Alleging that AI models can reproduce some works verbatim is probably the stronger argument, but AFAIK you have to coax them pretty hard to do so, and therefore AI companies might be able to argue they're tools like photocopiers or such. They can try that argument but it'll fall flat when you consider that a photocopier is reproduction agnostic, while LLMs generally have a ton of work going into them to prevent them from outputting damaging things (and they still fail). That fact makes them not at all comparable to a photocopier, setting aside the more obvious "subscription software service" different. Also, you "know" pretty wrong about the effort required. For a recent example, see: https://www.latimes.com/entertainment-arts/business/story/2024-06-24/riaa-suno-udio-lawsuit-ai-copyright-songs-music https://www.latimes.com/entertainment-arts/business/story/20... Here a number of people noticed getting specific producer tags basically unaltered in the output when just asking for songs of a certain genre, which then also often sound similar to existing songs.
- edent 2y agoBecause lots of AI firms literally pirated books to train their model. See https://shkspr.mobi/blog/2023/07/fruit-of-the-poisonous-llama/ https://shkspr.mobi/blog/2023/07/fruit-of-the-poisonous-llam... Neither the authors nor publishers received any compensation for having their work ingested. It isn't like OpenAI went to Amazon and bought one copy of every book - they downloaded a torrent.
- xdennis 2y agoI find English-law's way of legislating through the judiciary terrible. When a new situation appears (like whether machines should be allowed to learn on copyrighted works), the legality of the situation should be decided explicitly by the legislative branch, not by arcane interpretations of previous judicial decisions.
- jcranmer 2y ago> When a new situation appears (like whether machines should be allowed to learn on copyrighted works), the legality of the situation should be decided explicitly by the legislative branch Okay, so what happens in the interim situation? If the legislature hasn't spoken yet, is it assumed to be legal or assumed to be illegal? Or is this assumption tested on a case-to-case basis, with both sides making arguments as to why it should be treated to be legal/illegal in this specific scenario?
- jcranmer 2y agoFair use depends very heavily on the nature of the use. A key part of Google's defense was that not only was it not using the entire books to reproduce the entire book, but also that it was taking measures to prevent people from abusing Google's systems to reproduce an entire book. It's a lot of work to emphasis that the impact on the market (in other words, the fourth factor) is as minimal as practicable--and that's the crux of the analysis. When you're instead scanning someone's stock image database to build a tool to generate stock images... the fourth factor is jumping up and down screaming at you "YOU LOSE" and your best defense is that it's not the training, it's the tool built on the training data that is infringing the copyright.
- jiggawatts 2y agoLibraries would be illegal without centuries of precedent giving them a “legal inertia”. Google benefited from the exact same kind of bulk copyrighted data collection. They made verbatim copies of the text of both web sites and just about every book in existence! This kind of argument seems disingenuous to me. Either ban Internet search or acknowledge that training an AI on copyrighted text is no different than a student reading every book in a public library. Speaking of which: We all have free access to GPT 4o without advertising. It feels like asking a knowledgeable librarian.
- immibis 2y agoActually Google's faced this problem before. Caching is generally legal, and Google's just holding a cache to search through instead of requesting every site every time anyone searches.
- jiggawatts 2y agoThey show snippets from web sites and show subsets of books as well, including artworks and other diagrams in their entirety. For comparison, I had to fight for a year to get copyright permission to show book cover artwork in a library enquiry system! If I simply Google the same book titles or ISBNs, Google will show me the pictures directly. E.g.: https://www.google.com/search?q=greg+egan+eon&udm=2 https://www.google.com/search?q=greg+egan+eon&udm=2 How is that legal!? We had to pay to get access! In public and school libraries! The law in most western countries is very clear that book covers are "entire" works of art, and can only be displayed by organisations that pay the copyright holders. Google, Bing, and others violate copyright on a mass scale on a daily basis. Not to mention YouTube, TikTok, and Reels, all of which are packed wall-to-wall with "movie clips" and "TV show highlights". They're publishing copyrighted content uploaded by random people and then distributing the advertising revenue to the copyright violators instead of the copyright holders. This isn't "caching" or "indexing", it's verbatim serving.
- immibis 2y agoAnd they've faced legal battles over how much information they copy and display to users - not over the fact they do it at all.