3 ms·
> In principle, an AI learning from a scientific textbook is no different than a human student doing the same. It's not about the process of building a model.
by evdubs 2y ago
> In principle, an AI learning from a scientific textbook is no different than a human student doing the same.
It's not about the process of building a model. It's about the copyright infringement to acquire the training set and it's about the copyright infringement when producing output. Humans, just like these AI companies, can infringe copyrights by downloading copyrighted material. Humans, just like AI programs, can produce material that is derivative of material protected by copyright. Whatever the model represents and whatever is in your brain is not high on the list of concerns.
- worstspotgain 2y agoI've addressed the "when producing output" part in another reply, so here are my thoughts on the "to acquire the training set" part. Unless you're downloading from something like Z-Library (which incidentally is a deliberate attack on Western IP on Putin's behalf) you generally have the fair use rights to "digest" the information you receive. You can process the contents of a webpage or momentarily OCR a printed page, provided that it was legally distributed and that you don't store a copy proper. There can be specialized restrictions if it's something that you have a license for instead of a copy of, see e.g. https://en.wikipedia.org/wiki/Shrinkwrap_(contract_law) https://en.wikipedia.org/wiki/Shrinkwrap_(contract_law) As for derivative works, not all the works that are made "in proximity of" previous works are derivative works. You can make a simulated reality movie without infringing on the Matrix, for instance. Just write a somewhat different plot and don't name your lead character Neo.
- evdubs 2y ago> You can process the contents of a webpage or momentarily OCR a printed page, provided that it was legally distributed and that you don't store a copy proper. The AI companies are definitely storing proper copies of their "corpus". > You can make a simulated reality movie without infringing on the Matrix, for instance. You can also make a simulated reality movie, even without the name Neo, and a court could find it substantially similar to the Matrix and thus copyright infringing. It doesn't need to be the same.
- worstspotgain 2y ago> The AI companies are definitely storing proper copies of their "corpus". Their own lawyers would be on their case if they were infringing halfway through the pipeline. They don't have to, and it would weaken their case where it counts. > a court could find it substantially similar to the Matrix and thus copyright infringing Substantially similar is not the standard, otherwise they would have had to stop making time travel movies about 70 years ago.
- evdubs 2y agoSubstantially similar is the standard https://crsreports.congress.gov/product/pdf/LSB/LSB10922 https://crsreports.congress.gov/product/pdf/LSB/LSB10922 > Under U.S. case law, copyright owners may be able to show that such outputs infringe their copyrights if the AI program both (1) had access to their works and (2) created “substantially similar” outputs From Shaw v Lindheim. > Their own lawyers would be on their case if they were infringing halfway through the pipeline. This is why many of these AI companies are their own entities with sponsorships or shares held by other, larger companies. If they're on the hook for obscene copyright infringement, they just close down.
- worstspotgain 2y ago> Substantially similar is the standard This quote is about AI, not movies. If you read the entire report, there are many other caveats and provisions: Whether or not copying constitutes fair use depends on four statutory factors under 17 U.S.C. § 107: 1. the purpose and character of the use, including whether such use is of a commercial nature or is for nonprofit educational purposes; 2. the nature of the copyrighted work; 3. the amount and substantiality of the portion used in relation to the copyrighted work as a whole; and 4. the effect of the use upon the potential market for or value of the copyrighted work. and furthermore: The substantial similarity test is difficult to define and varies across U.S. courts.
- tensor 2y ago> It's about the copyright infringement to acquire the training set and it's about the copyright infringement when producing output. I just read your comment. I also read news articles I paid for today. None of that is copyright infringement, and scraping that same content to train an AI is no different. Yes, when an AI produces a substantially close reproduction that would be copyright infringement, but some systems have guards to prevent that, like Github Copilot.
- yencabulator 2y agoSmall scale personal use can be fair use while large-scale commercial use is not.
- evdubs 2y ago> None of that is copyright infringement, and scraping that same content to train an AI is no different. Scraping is copyright infringement. Try telling the RIAA that scraping songs from YouTube or Spotify is not copyright infringement. Try telling the MPAA that scraping movies from Netflix is not copyright infringement. Same for book publishers.