7 ms·
This settlement has basically nothing to do with LLMs. At least not as far as the courts are concerned. Alsup ruled [0] that feeding a book into an LLM is tran
by phire 3mo ago
This settlement has basically nothing to do with LLMs.
At least not as far as the courts are concerned. Alsup ruled [0] that feeding a book into an LLM is transformative and counts as fair use. Especially when they purchased a physical copy of the book, scanned it, and destroyed the original.
But if I'm reading the ruling correctly, Anthropic might have been fine even with feeding pirated books into their LLM (as long as they planned to eventually deleted them afterwards)
Where Anthropic ran into problems is that they put all their pirated books into a big central library (file on a server), and planned to keep those copies forever. Including copies they never actually fed into the LLM (a point that seriously worked against them).
Alsup ruled this central library of pirated books was copyright infringement. And it's this "pirated central library" that Anthropic are now paying a a 1.5B settlement for, nothing else.
The fact that the pirated books were also used to train LLMs is legally irrelevant. Though... I suspect a non AI company could have negotiated a significantly smaller settlement.
[0] https://copyrightalliance.org/wp-content/uploads/2025/06/Bartz-v.-Anthropic-Order.pdf https://copyrightalliance.org/wp-content/uploads/2025/06/Bar...
- embedding-shape 3mo ago> At least not as far as the courts are concerned. Alsup ruled [0] that feeding a book into an LLM is transformative and counts as fair use. Especially when they purchased a physical copy of the book, scanned it, and destroyed the original. > But if I'm reading the ruling correctly, Anthropic might have been fine even with feeding pirated books into their LLM (as long as they eventually deleted them afterwards) The way I understood it, was that essentially the entire case rested on if Anthropics use was "transformative" or not. And since they literally destroyed the books (not just delete files, which would be copied), that made it transformative. Regardless if they deleted files or not, if nothing existing was transformed, it would have been illegal. But because of the destruction of k̶n̶o̶w̶l̶e̶d̶g̶e̶ physical property, this ended up being legal.
- phire 3mo ago> And since they literally destroyed the books (not just delete files, which would be copied), that made it transformative. You have to be careful, just because the judge points a factor out as notable, doesn't mean that factor was required. The destruction of source books makes Anthropic's fair use argument [2] especially air tight, but it would be a mistake to assume that act was required, or is what made it transformative. In the previous google books case [1] (which this case cites), google borrowed books from libraries, scanned them, then returned them. They were not destroyed, google didn't even keep the physical copy. Yet Google Books was ruled fair use, because it was transformative. [1] https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,_Inc https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,.... [2] Note... This part of the ruling is still not about LLMs. This was about Anthropic's right to scan books and then keep a digital library of them.
- embedding-shape 3mo agoMy understanding comes from here, seems pretty clear to me but won't claim to be a lawyer of course: > Ultimately, Judge William Alsup ruled that this destructive scanning operation qualified as fair use—but only because Anthropic had legally purchased the books first, destroyed each print copy after scanning, and kept the digital files internally rather than distributing them. The judge compared the process to “conserv[ing] space” through format conversion and found it transformative. Had Anthropic stuck to this approach from the beginning, it might have achieved the first legally sanctioned case of AI fair use. Instead, the company’s earlier piracy undermined its position. https://arstechnica.com/ai/2025/06/anthropic-destroyed-millions-of-print-books-to-build-its-ai-models/ https://arstechnica.com/ai/2025/06/anthropic-destroyed-milli... Based on that I get the impression it's quite literally the destruction part that makes it transformative, without it, it wouldn't have been tranformative at all.
- phire 3mo agoI've read through the order again. I can't find anywhere where Alsup says the destruction was required. He cites three cases where a conversion from one format to another (without destruction of the previous version) was ruled to be fair use. Including scanning books with the google books case. (And referenced the Napster case, where a similar argument was rejected) Then made the following comparison. "Here, every purchased print copy was copied in order to save storage space and to enable searchability as a digital copy. The print original was destroyed. One replaced the other. And, there is no evidence that the new, digital copy was shown, shared, or sold outside the company. This use was even more clearly transformative than those in Texaco, Google, and Sony Betamax (where the number of copies went up by at least one), and, of course, more transformative than those uses rejected in Napster (where the number went up by “millions” of copies shared for free with others)." So it wasn't transformative because of the destruction. The destruction only made it "even more clearly transformative" than those other cases. Like, how can destruction be required if there were previous cases where it wasn't? The key legal point is not that Anthropic destroyed the books, but the key fact was that Anthropic didn't distribute the scanned copies. Alsup keeps returning to this point: "But what matters most is whether the format change exploits anything the Copyright Act reserves to the copyright owner. Anthropic already had purchased permanent library copies (print ones). It did not create new copies to share or sell outside" "But again, the replacement copy here was kept in the central library, not distributed" The conclusion of that section doesn't even mention the destruction at all. arstechnica isn't exactly wrong, the quote also mentioned "and kept the digital files internally rather than distributing them". It just put way too much emphasis on the destruction, and not enough on the lack of distribution. The other thing that arstechnica are missing: Antropic didn't destroy the books because they thought it would strengthen their legal argument. They destroyed the because it's a lot cheaper and faster to scan books by ripping off their bindings and feeding the stacks of loose pages into a document scanner.
- rpdillon 3mo agoYou're mistaken. The transformativeness of the use is independent of the destruction of the books. The destruction of the books allowed them to argue that they had not duplicated them, and was instrumental in the argument supporting the legality of scanning them. But that's entirely upstream of the way the data was leveraged, which is what is critical in the argument about the use being transformative.
- lp4v4n 3mo agoIt's easier to ask for forgiveness than permission, right? It seems to be the modus operandi of corporations in general: they commit any kind of infringement they want and then later they go for a settlement with a value that's, of course, not too big for a company too big to fail. In the meantime, the average person or company gets shafted. In my opinion, we are one step away from AI companies capturing the entirety of copyright legislation.
- baranul 3mo agoYou do indeed appear to have a valid point. Many "chosen" companies, like Uber for example, appear to have broken numerous laws. Legal action against many such companies comes suspiciously slowly, where they have already obtained massive profits and value, before the possibility of being shut down comes. Then, when they are finally pulled into court, they have all kinds of money for the best lawyers and have already paid the right politicians (and others). When the legal judgements for wrongdoing are finally handed out, they often come across as just an inconvenience or kind of tax, which is easily handled in comparison to the profits they've already made. Yet, if average Joe or persons not considered as being of "the right type" were to do such actions, they quickly get the full book thrown at them. Often, the full measure of legal punishment, where their company and life is or about nearly over.
- satvikpendem 3mo agoPaying this sort of fee in the first place is itself regulatory capture because only the big companies will be able to pay it. If they can pirate to make an LLM then so should us commoners be able to too.
- smolder 3mo agoAs far as I'm concerned, the courts are wrong, and training on ill gotten copyrighted material is not fair use. Given the clear value of highly trained LLMs, the investment they have taken on, and the amount of disruption to the existing economy they stand to make, in a just world, the people who created the training data deserve some level of compensation. I think, in the US, they are very afraid of falling behind China, who doesn't give a shit about intellectual property, but that doesn't mean we aren't crossing an ethical boundary, acting like them.
- davidguetta 3mo agocopyright is bullshit
- aetherspawn 3mo agoCopyright is what stops someone from copy+pasting a book that took years to write, then selling it $1 cheaper than the original author on Amazon or whatever and making a margin 1 million percent higher than the original author. Imagine a society without copyright… only physically intensive jobs could make money because everything else would be pirated, ripped-off or free. Thus, only those who are financially independent could afford to publish. Because the world really needs more rich class propaganda…
- smolder 3mo agoRight, it's about incentivising intellectual work. While I have big issues with the copyright system, like all the extensions lobbied for by Disney and friends, it did enable a lot of good work to happen.
- skinfaxi 3mo ago> it did enable a lot of good work to happen. How do we know that when we don't have a copy of the world without this regime? How much more and greater works could have been produced without such a repressive system? A really successful work becomes part of the culture, and remixing, derivatives and other modes of integrating cultural artifacts are prohibited. Why should we allow corporations to own our culture?
- satvikpendem 3mo agoThat's a good ruling, because otherwise only the big companies can afford to pay for enough content to make an LLM (say goodbye to open weight or research LLMs). Having a fee like this is actually a form of regulatory capture.
- cataphract 3mo ago> But if I'm reading the ruling correctly, Anthropic might have been fine even with feeding pirated books into their LLM (as long as they planned to eventually deleted them afterwards) The court says otherwise. > Such piracy of otherwise available copies is inherently, irredeemably infringing even if the pirated copies are immediately used for the transformative use and immediately discarded. Then it says it doesn't need to decide on that basis because they kept it not just for training LLMs, but also for building a central library. Which seems a bit ridiculous, because the sole purpose of the central library is to train LLMs.
- chirau 3mo ago[flagged]
- zelphirkalt 3mo agoWho would have thought, that this is the way, which we take to arrive at the burning books stage again? They neatly line up with historical perpetrators in that regard.
- freejazz 3mo ago>But if I'm reading the ruling correctly, Anthropic might have been fine even with feeding pirated books into their LLM (as long as they planned to eventually deleted them afterwards) You're not. Even if training is fair use, it doesn't mean you can steal copies to train the model. It just means the training itself isn't an infringement (in Alsup's opinion). Stealing the copies of the books was an infringement and that's exactly the liability that Anthropic settled.
- xuhu 3mo agoHey I just came up with this idea, I'm going to feed copyrighted books into my LLM that remembers them verbatim, and then people pay me to ask the LLM for complete copies of a book. Wait, no, not verbatim. It transforms upper case into lower case and vice versa.