9 ms·
More writers sue OpenAI for copyright infringement over AI training
- racked 3y agoIt's akin to suing a person for memorizing things from a book. Don't complain, go write something.
- buildbot 3y agoDo cliff notes of books and plays infringe/need a license? If so, that seems like they’d have a possible case. Maybe. If not, well… maybe openAI infringed by not buying their original copy, but not sure that feeding it into a bunch of math is going to be copyright infringement.
- throwaway290 3y ago> the system can accurately summarize their works and generate text that mimics their styles Cliff notes is not what lets you replicate the style of the author etc. And yeah, you can use "it's just feeding it into a bunch of math" to justify nearly anything that involves software including good old piracy. What matters is what math is used for. (Spoiler: line up Microsoft's pockets at the expense of actual writers in this case.)
- rmbyrro 3y agoNot at all like piracy. When someone pirates a book, they're replacing the original without consent or remuneration to the copyright holders. When you train an AI on the contents of a book, you're not replacing it. If someone is interested in the content, they still need to buy it. Using ChatGPT is not a substitute. If it is, they're gonna have to prove it in court, but I doubt they'll be able to.
- throwaway290 3y agoIf you can ask ChatGPT about any book contents, you don't need to get the book, and if you don't need to get the book then author got robbed, ClosedAI/MS profited.
- jnovek 3y agoThis reminds me of the tenuous RIAA claim that every pirated piece of media represented a lost sale back when they were suing their customers in the 2000s.
- tpmoney 3y agoIt's been something of a wild ride for me having lived through the "Information wants to be free" era to now live in the new "Reading my publicly published writings and deriving new things from that is theft" era. The next few years of court battles around this are going to be interesting, and I'm not too hopeful on the odds that the "little guy" wins in the end. Seemingly "little guy" affirming results might just turn around and further entrench large players instead.
- throwaway290 3y ago> Reading my publicly published writings and deriving new things from that is theft Ah, the mental gymnastics people go through to justify the theft. Just... no. It's nothing about people reading your writings and deriving things from that. It's about big companies using automated tools to ingest your writing and provide commercial services based on it. To other people. Without paying you a dime.
- true_religion 3y agoThe ability to ask people what is contained within a book isn't obviously copyright infringement. Merely summarizing info and attributing it to the source is the basic element of learning, for both machines and human beings. These suits are necessary becsuse it's not clear where the line is, and if ChatGPTs functions actually cross it. What is clear is that OpenAI is doing its best to avoid infringing anyone's copyright even if it is trivial for them to do so. They have the training data so they can simply output it word for word bypass the LLM. They don't do that and further restrain their LLM from making too long recitations. If you can trick / manipulate the LLM into giving you too much then I say that infringement is on you.
- extragood 3y agoCopyright doesn't protect style or genre [1], so these suits seem destined to fail. That said, it seems like it is time to reexamine those laws in light of current technology before it kills off creative works. https://creativecommons.org/2023/03/23/the-complex-world-of-style-copyright-and-generative-ai/#:~:text=Copyright%20doesn't%20protect%20things,express%20themselves%20through%20their%20works https://creativecommons.org/2023/03/23/the-complex-world-of-....
- AlbertCory 3y agoI'm miffed. I tried a couple characters from my books, and zilch: ===== who is dan markunas ChatGPT I'm sorry, but I don't have any information on a person named Dan Markunas in my database .... who is janet saunders ChatGPT I'm sorry, but I don't have any specific information about a person named Janet Saunders in my database, ===========
- SanderNL 3y agoYour book was published somewhere mid-2021, right?
- AlbertCory 3y agoTwo books. The second was after their cutoff date.
- probably_wrong 3y agoRandom thought: my blog is licensed under a Creative Commons license [1] that allows you to use and transform my content as long as you give attribution and distribute your contributions under the same terms. I found the OpenAI bot scraping my blog recently. Assuming they used that data, when will they attribute me? [1] https://creativecommons.org/licenses/by-sa/4.0/ https://creativecommons.org/licenses/by-sa/4.0/
- guiambros 3y ago> Assuming they used that data... That's the key part. You haven't yet proved they have actually used your content for anything (other than, potentially, read the license to decide if they should include or discard from their training set). But in practice we'll never know for sure if they are respecting the terms of licenses until 1) this is tested in court, or 2) there's some internal leak that points into either direction.
- JamisonM 3y agoI expect that OpenAI would concede that they used the data in any court case immediately to get that issue off the table, I really don't think they have a strong interest in foot-dragging on this stuff, right? I would think OpenAI wants the thornier legal issues actually settled so that the whole ecosystem can grow within those terms & they can lobby for the legal changes they need/want?
- lokar 3y agoThe alternative would be discovery on that issue, which they may want to avoid.
- mistrial9 3y ago> wants the thornier legal issues actually settled .. wants the thornier issues to be debated and re-tried ad infinitum, as long as they generate cash flow and build their moat(s).. more likely
- mistrial9 3y agothis is a great lawsuit! if you read the complaint, they catch OpenAI dead-to-rights .. asking about plot details with names from the books, asking to write a paragraph in the style of the author in that book, and a diversity of authors that shows social awareness.. great support for this from California
- tpmoney 3y agoIf you go to wikipedia and look up a book, you'll likely find plenty of plot details, including character names. Is this also infringing? As far as style goes, copyright doesn't protect that. Trademark MIGHT if your style is distinctive enough to be a trademark (and is used as such), but the "style" of a writer is largely about tempo and word choices, none of which are subject to copyright protections.
- mistrial9 3y agoI think we are now reproducing multiple generations of debate on this topic, in a few go-rounds.. Let's note that among the four largest economies in the world, they each have different rules for this.
- tpmoney 3y agoDo any of those economies really have laws protecting an author's "style"? Because I'd really like to see a legal definition of an author's style, and a case that found someone guilty of infringing on that style (separate from trademark and copyright of course)
- tsegratis 3y agoI'm waiting for the day OpenAI sues humans for infringement of it's prompt output
- mannyv 3y agoAll these lawsuits will die. Why? Because people train on corpuses of data all the time, without a license or any attribution. Every piece of text a writer reads is training that writer. Every image an artist sees helps to train that artist. Every sound a musician hears is training that musician. That doesn't mean they can't exclude their works from training via a license going foreward. But that becomes an enforcement problem.
- plagiarist 3y agoIIRC courts have already ruled AI-generated works cannot have copyright. So there is already a legal distinction between a human and a model creating works. I also doubt "humans are just a larger Markov chain than the LLM and they're allowed to" will hold up in court.
- brookst 3y agoI don’t see what eligibility to have works protected has to do with legality of learning. I really hope “copyright can be used to prohibit reading and learning” does not hold up in court. Copyright is, and should be, a protection from unauthorized reproduction. Extending it to protect the abstract ideas would be a disaster. And extending it to control stylistic learning would be even worse.
- plagiarist 3y agoYou are not understanding and making a lot of assumptions isn't a substitute for that.
- brookst 3y agoHard to argue with that level of reasoning and sourcing.
- anonymousab 3y agoPeople can do a lot of things that we don't legally allow machines or automation to do.
- yieldcrv 3y agoThe fair use argument is quite strong If you dissect the plaintiffs claim they are arbitrarily conflating training and regurgitating Training is using for criticism and comparison purposes, hence fair use And there is no lawsuit against what it regurgitates and the purpose of its output, whether someone asks it to give a list for comparison purposes, or specifically asks it for a story that has a plagiarized result
- RecycledEle 3y agoIs there a list of the critters suing AI companies so I can boycott them?
- lamp987 3y agoWhy would you want to do that?
- RecycledEle 3y agoAI is the next industrial revolution. It will greatly increase human productivity. Anyone standing against it is an enemy of all mankind.
- hnis4fascists 3y ago[flagged]