3 ms·
Copying into RAM during training is making a copy, and can be a copyright violation. https://en.wikipedia.org/wiki/MAI_Systems_Corp._v._Peak_Computer,_Inc http
by jpollock 3y ago
Copying into RAM during training is making a copy, and can be a copyright violation.
https://en.wikipedia.org/wiki/MAI_Systems_Corp._v._Peak_Computer,_Inc https://en.wikipedia.org/wiki/MAI_Systems_Corp._v._Peak_Comp....
However, it seems that there is a later case in the 2nd circuit:
https://en.wikipedia.org/wiki/Cartoon_Network,_LP_v._CSC_Holdings,_Inc https://en.wikipedia.org/wiki/Cartoon_Network,_LP_v._CSC_Hol....
- koboll 3y agoThe difference is that the copy is authorized, unless the work is being pirated. When an artist displays their work on DeviantArt or Artstation or whatever, they are allowing the general public to load it into memory. It's part of the license agreement they sign when they sign up for these services.
- jpollock 3y agoThe copy isn't authorized, the copy is allowed under Fair Use. There's a huge difference between the two.
- koboll 3y agoWrong. Fair Use applies to instances that would otherwise be copyright violations, i.e. unauthorized distribution. When you sign up for a social media site you EXPLICITLY grant the site the rights to distribute it. You have expressly permitted it. It's a big difference!
- nearbuy 3y agoThe sources used for training these AIs are publicly available sources like Common Crawl. If having a copy in RAM is a copyright violation, then there are copyright violations occurring well before any AI ever sees it.
- emkoemko 3y agoit is and the same reason Blizzard can sue cheat makers because they are violating copyright law by using the memory of the game etc
- fnordpiglet 3y agoHow do search engines exist? The internet archive? Caching of image results? Web browser caches? CDNs?
- jpollock 3y agoCopies made by search engines don't need authorization, and can be unauthorized copies. Search engines are allowed to make copies under Fair Use since they are transformative - see Authors Guild, Inc. v. Google, Inc. There hasn't been an explicit decision for ML training, but everyone's assuming that Authors Guild v Google applies. https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,_Inc https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,..... CDNs operate under the control of the copyright owner, so they would be authorized. Web browser caches are under the control of the recipient who has authorization to make a copy.
- harshreality 3y agoMAI v. Peak was obviously wrong. It would mean whenever you use someone else's computer, and run licensed software, you're committing copyright infringement. The decision split hairs distinguishing between the current user and the licensee for purposes of legality of making transient copies in memory as part of running the program. Peak was a repair business. MAI built computers (as in assembled/integrated; I think they were PCs) and had packaged an OS and some software presumably written or modified in-house along with the computer. MAI serviced the whole thing as a unit. So did Peak. MAI sued Peak for copyright infringement because Peak was taking computer repair/maintenance business away from MAI, under the theory that Peak employees operating their clients' MAI computers and software was copyright infringement. (There were other allegations of Peak having unlicensed copies of MAI's software internally, but that's not central to the lawsuit.) If you have a piece of IP to use to train an IP model with, and you have legal right of access to use that piece of IP (for private purposes), MAI v. Peak doesn't cleanly apply. MAI v. Peak is also 9th circuit only, and even without the poor reasoning, it should automatically be in doubt because the 9th circuit is notoriously friendly to IP interests, given that it covers Los Angeles.
- jpollock 3y agoI agree that MAI v Peek is crazy. I was only pointing out that the law is of the opinion that a copy is a copy is a copy, regardless of where it's made, or how long it exists for. Other decisions come into play to save us, like Authors Guild v Google, where they said search engines could make copies, bringing Fair Use into the picture. Personally, I think that creating the model is Fair Use, but anything produced by the model would need to be checked for a violation. I would treat it the same as if I went to Google Book Search, and copied the snippet it returned into my new book. The license associated with the training data then becomes insanely important. Having the model reference back to the source data is even more important. For example, training data with a CC BY license would be very different to CC BY-SA and CC BY-ND, and they all require the work produced by the model to have credit back to the original source to be publishable. https://creativecommons.org/licenses/ https://creativecommons.org/licenses/