6 ms·
People keep comparing it to a human ingesting content throughout their life and then being influenced in their own works. I'm sorry but that is not the same thi
by everly 3y ago
People keep comparing it to a human ingesting content throughout their life and then being influenced in their own works. I'm sorry but that is not the same thing - at all.
The concept of training a model with the explicit intent of selling the output of that model is inherently different.
Not saying that it should be illegal. But it is clearly in violation of the spirit of existing copyright law, in my opinion.
They set out with the intent to make money, using copyrighted input. Seems pretty simple.
See dragonwriter's comment for the articulate version of what I'm saying
- artninja1988 3y agoIt is not illegal to make money with copyrighted content. Nor should it be
- everly 3y agoIf you don't procure the appropriate license for that content then yes, it is (exceptions for things like parody notwithstanding). For example, clearance of music samples.
- nikanj 3y agoSo aspiring authors should avoid all reading, because they need licenses for all works that inspired them?
- everly 3y agoThat is an absurd conclusion to draw. Authors do not need licenses for works that inspired them.
- areoform 3y agoThen why should AI?
- saulpw 3y agoBecause AI isn't human? It's a machine/algorithm?
- akasakahakada 3y agoThen why are these exceptions unique to human?
- saulpw 3y agoTake it up with the US Copyright Office. Per GPT4: In the "monkey selfie" case, a photographer named David Slater set up a camera in the Indonesian jungle, and a macaque monkey took a photograph of itself with it. When the photo was uploaded and shared, various parties began to argue over who held the copyright. Slater claimed it was his because it was his camera and he set up the situation. Others believed that if the monkey pressed the shutter, then the monkey, or no one, held the copyright. The U.S. Copyright Office clarified its stance on the matter in the Compendium of U.S. Copyright Office Practices, Third Edition. It stated: "The U.S. Copyright Office will not register works produced by nature, animals, or plants. Likewise, the Office cannot register a work purportedly created by divine or supernatural beings, although the Office may register a work where the application or the deposit copy state that the work was inspired by a divine spirit."
- ipython 3y agoBut did you not purchase that content first?
- chihuahua 3y agoBorrowed it from the library.
- melagonster 3y agobut buying a book is not buying the copyright of it.
- metalspot 3y agoThe principal criminal statute protecting copyrighted works is 17 U.S.C. § 506(a), which provides that "[a]ny person who infringes a copyright willfully and for purposes of commercial advantage or private financial gain" shall be punished as provided in 18 U.S.C. § 2319. Section 2319 provides, in pertinent part, that a 5-year felony shall apply if the offense "consists of the reproduction or distribution, during any 180-day period, of at least 10 copies or phonorecords, of 1 or more copyrighted works, with a retail value of more than $2,500." 18 U.S.C. § 2319(b)(1). https://www.justice.gov/archives/jm/criminal-resource-manual-1847-criminal-copyright-infringement-17-usc-506a-and-18-usc-2319 https://www.justice.gov/archives/jm/criminal-resource-manual...
- imgabe 3y agoYeah and people read copyrighted books so they can get jobs and sell the use of the knowledge they acquired from reading the books. The only difference is a machine doing it at a larger scale.
- MegaDeKay 3y agoAnd (most) people buy those copyrighted books.
- imgabe 3y agoLibraries exist. Borrowing books from a friend exists. Even if you stole the book, using the knowledge you learned from it to make money is not copyright infringement. If you steal a book about how to make money flipping houses and then start flipping houses, at worst you are on the hook for a minor theft. The author can't sue you for illicitly learning to flip houses from the stolen book. That is just not how that works at all.
- data-ottawa 3y agoThe library is allowed to lend out books, they're not producing copies. I can lend you a book without producing a copy. I can steal a book without making a copy. I can learn from a book without making a copy. ChatGPT was trained by copying works into a dataset and using that for commercial purposes. Surely you can see how that's different.
- everly 3y agoThose people are entering into and abiding by a license agreement when they purchase the book.
- imgabe 3y agoNo they aren't. Where's the Terms of Service for a book? What about secondhand books or books you get from the library or books you borrow from a friend? I never had a book tell me to an accept a license agreement before I could read it.
- mdorazio 3y agoI'm not understanding how it's different. For example, if I specifically set out to make money by creating and selling a parody of Harry Potter by reading all the Harry Potter books a bunch of times, does that make my parody a violation of the spirit of existing copyright law? Edit to add another example because someone is going to say that parody is its own thing and exempted. If I want to make money by writing a film in the same way Tarantino does I go and read all his screenplays to understand his style, pacing, character archetypes, etc. Should that also not be ok?
- everly 3y agoNo, because parody is a specifically protected category under the law [0]. But, as with many things, the legality would ultimately depend on the nuances of the implementation. [0] https://www.law.cornell.edu/wex/parody https://www.law.cornell.edu/wex/parody
- evdubs 3y agoDid you, through copyright infringement, acquire your copies of Harry Potter? Did you use Harry Potter source material and a machine to transform the input and produce your parody material? I think it's different if OpenAI acquires licenses for all of the material it uses for training.
- mdorazio 3y agoI sure didn't buy a license for every book and screenplay I've ever read. And I borrow a lot of books from the library. I'm not understanding why you think ML models need to license all of their input content when humans clearly do not.
- evdubs 3y agoIf you bought the books and screenplays, you acquired them in a way that doesn't infringe on copyright. If you borrowed them from the library, you acquired them in a way that doesn't infringe on copyright. You don't need a license because you aren't republishing. ML training doesn't work without having a copy of the data (books and screenplays). That data can either be copied in a non-infringing way (buy the books; acquire a license) or in an infringing way (download the books from a corpus without explicit permission like these AI companies are accused of doing). LLM services need a license just like Facebook, Instagram, etc. need a license to republish the stuff you post. The copyright holder maintains their copyright and the service that republishes their work (Facebook, Instagram, etc.) or publishes derivative works (AI) needs a license.
- brianjking 3y agoPlease explain how this is any different than the Google Books case. Please do so without any feelings involved to the best of your ability. Even if OpenAI maliciously ingested copywritten work, it is just a bunch of numbers and if you go in and ask for it spit the book back out, it won't. That simply isn't how this works.
- everly 3y agoI don't see much of a difference, seems like google profited off of copyrighted material that they didn't own. I dont understand the purpose of your feelings comment. I have no affiliations or preferences for any dogs in this fight. I don't necessarily think they are identical scenarios but if I were OpenAI's lawyer that's probably what I'd try and point at. I didnt say it would spit the book back out, at any point. I've said a human chose to copy the entirety of copyrighted texts, verbatim, into the training corpus for a language model that they intended to sell.
- brianjking 3y agoGoogle isn't selling any books, nor is it distributing the book. OpenAI isn't pirating any books content nor is it distributing or reproducing it. Nor are they even "stealing the book". Both are transformative content.
- everly 3y agoI don't agree that they aren't pirating content. If full text is used in the training corpus that seems like it would qualify. But I understand what you're saying and it's certainly possible you're correct, I just don't think it's obvious.
- jameshart 3y ago> The concept of training a model with the explicit intent of selling the output of that model is inherently different Uh-oh, better tell the colleges and universities to stop promoting degree programs off the back of the potential increase in lifetime earnings. Wouldn't want anyone to get the idea that it's okay to absorb a bunch of copyrighted material during college to update their mental model, then make money by selling the output of that model for the rest of their lives.
- 28304283409234 3y agoYou do understand that those are humans we are talking about. Humans with lives and livelihoods and wants and needs. Wishing to learn for the sake of learning and growing as well as sustaining themselves? And that openai is a company trying to make money? You do see the enormous difference there I hope? Let alone the huge shift of power from one to many (one university to many students), to many to one (many resources to one company).
- jameshart 3y agoI don’t get the ‘OpenAI is a company not a person’ argument. Would those arguments go away if the GPT model had been published by a single individual, Satoshi style? The question of whether it’s okay for OpenAI to train an LLM on copyrighted material and sell the results would also go to whether you or I can train an LLM (or finetune one, or do any other kind of ML training), on copyrighted material and make commercial use of the result. It seems to me OpenAI have done the equivalent of distilling a bunch of knowledge down into a book. And they’ve given that book a really good index. They’re essentially an encyclopedia vendor. Just an exceedingly sophisticated one. Back in the olden times publishers used to pay people to write encyclopedia entries based on summarizing stuff they had read in other books. Nowadays we still do the same thing only with volunteer time (Wikipedia requires everything it contains to be externally sourced, after all. You can’t write anything into a Wikipedia article without reading it somewhere copyrighted first). OpenAI’s model isn’t so different. Source material, indexed and summarized to make it easier to search and use.
- akasakahakada 3y agoYes, it is different that human has magic and machine don't. I learn stuff from books and sell my skills. I definitely copied someone's knowledge about calculus, engineering, and I am intented to earn money from them. Machine cannot do this because only me has magic.
- melagonster 3y agowhy computer in openAI have magic can let author lose their copyright ? you know, because they have a magical algorithm?
- akasakahakada 3y agoYes. Deep learning is not copy and paste at all, and this is the whole point. If you ask ChatGPT to quote stuff from a famous book, of course you will get what you want. It is prompt engineering. Human are able to quote stuff verbatim too. But just like you cut open the brain, you open up the Numpy array weight, all you can see is just nonsense, you don't see any bookshelf sitting in there with a pile of copyrighted books.
- melagonster 3y agoI don't know, this convert some book to unreadable format, but I still input these book into computer and convet it to another thing. maybe there are another good reason for permit this happening, but I just personally do not like people yelled "AI is human!" or "you just another larger matrix!". (sorry for passive aggressive)
- bathtub365 3y agoScale is an important distinguishing factor and also breaks the “human reads book and writes something inspired by it” analogy. A human can’t read all books ever written and then copy themselves an infinite number of times and create an infinite number of inspired works simultaneously but an ML model can.