9 ms·
Here's the full decision, which (like most decisions!) is largely written to be legible to non-lawyers: https://storage.courtlistener.com/recap/gov.uscourts.ded
by JackC 2y ago
Here's the full decision, which (like most decisions!) is largely written to be legible to non-lawyers: https://storage.courtlistener.com/recap/gov.uscourts.ded.72109/gov.uscourts.ded.72109.770.0.pdf https://storage.courtlistener.com/recap/gov.uscourts.ded.721...
The core story seems to be: Westlaw writes and owns headnotes that help lawyers find legal cases about a particular topic. Ross paid people to translate those headnotes into new text, trained an AI on the translations, and used those to make a model that helps lawyers find legal cases about a particular topic. In that specific instance the court says this plan isn't fair use. If it was fair use, one could presumably just pay people to translate headnotes directly and make a Westlaw competitor, since translating headnotes is cheaper than writing new ones. And conversely if it isn't fair use where's the harm (the court notes no copyright violation was necessary for interoperability for example) -- one can still pay people to write fresh headnotes from caselaw and create the same training set.
The court emphasizes "Because the AI landscape is changing rapidly, I note for readers that only non-generative AI is before me today." But I'm not sure "generative" is that meaningful a distinction here.
You can definitely see how AI companies will be hustling to distinguish this from "we trained on copyrighted documents, and made a general purpose AI, and then people paid to use our AI to compete with the people who owned the documents." It's not quite the same, the connection is less direct, but it's not totally different.
- echelon 2y agoIf the copyright holders win, the model giants will just license. This effectively kills open source, which can't afford to license and won't be able to sublicense training data. This is very bad for democratized access to and development of AI. The giants will probably want this. The giants were already purchasing legacy media content enterprises (Amazon and MGM, etc.), so this will probably further consolidation and create extreme barriers to entry. If I were OpenAI, I'd probably be very happy right now. If I were a recent batch YC AI company, I'd be mortified.
- dkjaudyeqooe 2y agoLicense what? Every available copyrighted work? Even getting a tiny fraction is not practical. To the contrary, this just means companies can't make money from these models. Those using models for research and personal use wouldn't be infringing under the fair use tests.
- echoangle 2y agoThey didn’t train it on every available copyrighted work though, but on a specific set of legal questions and answers. And they did try to license them, and only did the workaround after not getting a license.
- synthetic-echo 2y agoI think they were talking about the "model giants" like OpenAI you mentioned. Not saying they're correct, but I will concede the amount of copyrighted information someone like OpenAI would want is probably (at least) an order of magnitude more than this particular case.
- mvdtnz 2y ago> License what? Every available copyrighted work? Even getting a tiny fraction is not practical. Oh no. Anyway.
- marcus0x62 2y ago> License what? Every available copyrighted work? Even getting a tiny fraction is not practical. Maybe the strategy is something like this: 1) Survive long enough/get enough users that killing the generative AI industry is politically infeasible. 2) Negotiate a compromise similar to the compulsory mechanical royalty system used in the music business to “compensate” the rights holders whose content is used to train the models The biggest AI companies could even run the enforcement cartels ala BMI/ASCAP to compute and collect royalties owed. If you take this to its logical conclusion, the AI companies wouldn’t have to pre-license anything, and would just pay out all the royalties to the biggest rights holders (more or less what happens in the music industry) on the basis that figuring out what IP went into what model output is just too hard, so instead they just agree to distribute it to whomever is on the New York Times best seller list at any given moment.
- JoshTriplett 2y ago> If the copyright holders win, the model giants will just license. No, they won't. The biggest models want to train on literally every piece of human-written text ever written. You can pay to license small subsets of that at a time. You can't pay to license all of it. And some of it won't be available to license at all, at any price. If the copyright holders win, model trainers will have to pay attention to what they train on, rather than blithely ignoring licenses.
- simonw 2y ago"The biggest models want to train on literally every piece of human-written text ever written" They genuinely don't. There is a LOT of garbage text out there that they don't want. They want to train on every high quality piece of human-written text they can get their hands on (where the definition of "high quality" is a major piece of the secret sauce that makes some LLMs better than others), but that doesn't mean every piece of human-written text.
- ethbr1 2y agoEven restricted to that narrower definition, the major commercial model companies wouldn't be able to afford to license all their high-quality human text. OpenAI is Uber with a slightly less ethically despicable CEO. It knows it's flaunting the spirit of copyright law -- it's just hoping it could bootstrap quickly enough to make the question irrelevant. If every commercial AI company that couldn't prove training data provenance tomorrow was bankrupted, I wouldn't shed an ethical tear. Live by the sword, die by the sword.
- brookst 2y agoBold idea, requiring startups to proactively prove they have not broken the law. Should we apply it to all tech startups? Let’s see silicon startups prove they have not stolen trade secrets!
- vkou 2y agoI don't have either a data center, or every single copyrighted work in history to import as training data to train my open source model. Whether or not OpenAI is found to be breaking the law will be utterly irrelevant to actual open AI efforts.
- mvdtnz 2y agoOpen source model builders are no more entitled to rip off content owners than anyone else. I couldn't possibly care any less if this impacts "democratized access" to bullshit generators. At least if the big boys license the content then the rightful owners get paid (and have the option to opt out).
- brookst 2y agoThe copyright lobby has really done a number on public policy. Copyright was never meant to be perpetual. I’m good with your proposal if we also revert to the original 14 year + 14 year extension model. As it stands the 120 year copyright is so ridiculously tilted that we should not allow it to extend to veto power over technical advancements.
- Tanjreeve 2y agoLegal arbitrage isn't a technical advancement. The technical advancement was all the stuff that goes into LLMs not the part where we feed ever more copyright into models for AICorp to make money.
- pabs3 2y agoOpen source models can crowdsource open source training data. This was done for RNNoise for example.
- dkjaudyeqooe 2y ago> The court emphasizes "Because the AI landscape is changing rapidly, I note for readers that only non-generative AI is before me today." But I'm not sure "generative" is that meaningful a distinction here. Also the judge makes that statement, it looks like he misunderstands the nature of the AI system and the inherent generative elements it includes.
- deleted 2y ago[deleted]
- echoangle 2y agoHow is the system inherently generative?
- currymj 2y agoGenerative is a technical term, meaning that a system models a full joint probability distribution. For example, a classifier is a generative model if it models p(example, label) -- which is sufficient to also calculate p(label | example) if you want -- rather than just modeling p(label | example) alone. Similar example in translation: a generative translation model would model p(french sentence, english sentence) -- implicitly including a language model of p(french) and p(english) in addition to allowing translation p(english | french) and p(french | english). A non-generative translation model would, for instance, only model p(french | english). I don't exactly understand what this judge meant by "generative", it's presumably not the technical term.
- echoangle 2y agoDo you have some kind of dictionary where I can find this definition? Because I don’t really understand how that can be the deciding factor of „generative“, and the wiki page for „generative AI“ also seems to use the generic „AI that creates new stuff“ meaning. By your definition, basically every classifier with 2 inputs would be generative. If I have a classifier for the MNIST dataset and my inputs are the pixels of the image, does that make the classifier generative because the inputs aren’t independent from each other?
- Ajedi32 2y agoInterestingly, almost the entirety of the judge's opinion seems to be focused on the question of whether the translated notes are subject to copyright. It seems to completely ignore the question of whether training an AI on copyrighted material constitutes making a copy of that work in the first place. Am I missing something? The judge does note that no copyrighted material was distributed to users, because the AI doesn't output that information: > There is no factual dispute: Ross’s output to an end user does not include a West headnote. What matters is not “the amount and substantiality of the portion used in making a copy, but rather the amount and substantiality of what is thereby made accessible to a public for which it may serve as a competing substitute.” Authors Guild, 804 F.3d at 222 (internal quotation marks omitted). Because Ross did not make West headnotes available to the public, Ross benefits from factor three. But he only does so as part of an analysis of whether there's a valid fair use defense for Ross's copying of the head notes, ignoring the obvious (to me) point that if no copyrighted material was distributed to end users, how can this even be a violation of copyright in the first place?
- unyttigfjelltol 2y agoRoss evidently copied and used the text himself. It's like Ross creating an unauthorized volume of West's books, perhaps with a twist. Obscurity ≠ legal compliance.
- Ajedi32 2y agoSo the use of AI actually has nothing to do with the ruling here? This is just about the fact that Ross made one local copy of the notes and never distributed it?
- brookst 2y agoHow would training on copyrighted material be infringement in a way that merely producing the training material (but not iterating through training) would not be?
- anon373839 2y agoThis is an interesting opinion, but there are aspects of it that I doubt will stand the test of time. One aspect is the court’s ruling that West’s headnotes are copyrightable even when they merely quote a court opinion verbatim, because the editorial decision to quote the material itself shows a “creative spark”. It really isn’t workable — in law specifically - for copyright to attach to the mere selection of a quote from a case to represent that case’s holding on an issue. After all, we would expect many lawyers analyzing the case independently to converge on the same quotes! The key fact underlying all of this, I think, is that when Ross paid human annotators to write their own versions of the headnotes, they really did crib from West’s wholesale rather than doing their own independent analysis. Source text was paraphrased using curiously similar language to West’s paraphrasing. That, plus the fact that Ross was a directly competing product, is what I see as really driving this decision. The case has very little to say about the more commonly posed question of whether copyright is infringed in large-scale language modeling.
- zozbot234 2y agoIf close paraphrase can be detected, this ought to be proof enough that some non-trivial element of creativity was involved in the original text. Because purely functional and necessary elements are not protected by copyright, even when they would otherwise be creative (this is technically known as the 'scenes à faire' case) - and surely a "quote" which is unavoidable because it factually and unquestionably is the core of the ruling would have to fall under that.
- DrScientist 2y agoIsn't the argument that the act of selecting the right quote is the real work - and the work the copier avoided in the act of copying? You could argue that all the words are already in the dictionary - so none of them are new, you are just quoting from the dictionary in a particular order...... The reason you have people, rather than computers interpreting the law, is you can make judgements that make sense. Fundamentally these laws are there to protect work being unfairly ripped off. What was clearly done in this case was a rip-off which damaged the original creator - everything else is dancing on the head of a pin.
- reissbaker 2y agoI'll quote a longer portion of the transcript about generative AI, because I think it makes the opposite of your point: Ross’s use is not transformative. Transformativeness is about the purpose of the use. “If an original work and a secondary use share the same or highly similar purposes, and the second use is of a commercial nature, the first factor is likely to weigh against fair use, absent some other justification for copying.” Warhol, 598 U.S. at 532–33. It weighs against fair use here. Ross’s use is not transformative because it does not have a “further purpose or different character” from Thomson Reuters’s. Id. at 529. Ross was using Thomson Reuters’s headnotes as AI data to create a legal research tool to compete with Westlaw. It is undisputed that Ross’s AI is not generative AI (AI that writes new content itself). Rather, when a user enters a legal question, Ross spits back relevant judicial opinions that have already been written. D.I. 723 at 5. That process resembles how Westlaw uses headnotes and key numbers to return a list of cases with fitting headnotes. I think it's quite relevant that this was not generative AI: the reason that mattered is that "transformative" use biases towards Fair Use exemptions from copyright. However, this wasn't creating new content or giving people a new way to understand the data: it was just used in a search engine, much like Westlaw provided a legal search engine. The judge is pointing out that the exact implementation details of a search engine don't grant Fair Use. This doesn't make a ruling about generative AI, but I think it's a pretty meaningful distinction: writing new content seems much more "transformative" (in a literal sense: the old content is being used to create new content) than simply writing a similar search engine, albeit one with a better search algorithm.
- BoorishBears 2y agoI came here to point this out, and it's especially clear if you contextualize this with the original decision from September: https://www.ded.uscourts.gov/sites/ded/files/opinions/20-613_1.pdf https://www.ded.uscourts.gov/sites/ded/files/opinions/20-613... They were doing semantic search using embeddings/rerankers. The point that reading both decisions together compounds is that if they had trained a model on the Bulk Memos and generated novel text instead of doing direct searches, there likely would have been enough indirection introduced to prevent a summary judgement and this would have gone to a jury as the September decision states. In other words, from their comment: > But I'm not sure "generative" is that meaningful a distinction here. The judge would not seem to agree at all.
- kevin_thibedeau 2y agoThere were data brokers who literally paid people to transcribe phone books before OCR was a viable option. That was protected, as data isn't copyrightable. It isn't hard to argue that case law metadata is no different even though it includes textual descriptions (themselves taken from public documents).
- AlexCoventry 2y ago> You can definitely see how AI companies will be hustling to distinguish this from "we trained on copyrighted documents, and made a general purpose AI, and then people paid to use our AI to compete with the people who owned the documents." It's not quite the same, the connection is less direct, but it's not totally different. Surely creating a general-purpose AI is transformative, though? Are you anticipating that AI companies will be sued for contributory infringement, because customers are using a general-purpose AI to compete with companies which created parts of the training data?
- llamaimperative 2y agoIMO yes. The entire purpose of copyright law is to protect the incentive to create new material. A huge portion of the value prop of AI is that it captures the incentive normally bound for the creators of the training material (i.e. the whole point is you can ask the AI and not even see, never mind pay, the originator).
- AlexCoventry 2y agoI'm not a lawyer, but I think the bar for contributory infringement is much higher than that. I think you'd have to find representatives of the defendants actually indicating somehow that people should use it that way. It seems to me that Grokster, etc.'s encouragement of their users to infringe copyright was an important factor in them losing this case, for instance. https://supreme.justia.com/cases/federal/us/545/913/ https://supreme.justia.com/cases/federal/us/545/913/
- llamaimperative 2y agoEncouragement is definitely not a required element. https://www.cantorcolburn.com/news-newsletters-387.html https://www.cantorcolburn.com/news-newsletters-387.html
- zozbot234 2y agoAsk the AI for what exactly? Factual information? That gets very low protection from a copyright point of view, especially when separate random answers by the AI will routinely show completely different rephrasings of the AI's response - implying that it can generalize well beyond the "expression" contained in any single answer, and effectively reference the underlying facts.
- qingcharles 2y agoWestlaw's headnotes are primarily just snippets of the case with tags attached. They are really crappy. I hate them. Some lawyers love them. Westlaw protects them because they are the "value add." Otherwise their business model is "take published decisions the court is legally bound to provide for free and sell it to you." An LLM today could easily recreate the headnotes in a far superior manner from scratch with the right prompt. I don't even think hallucinations would factor in on such a small task that was well regulated, but you can always just asterisk the headnotes and put a disclaimer on them.
- Tteriffic 2y agoExactly. Why use the headnotes at all? I always thought they were obviously were copyrightable. Plus they’re not close to perfect either.
- alberto-m 2y ago> Here's the full decision, which (like most decisions!) is largely written to be legible to non-lawyers For me (Italian) this is amazing! Most Italian judges and lawyers write in a purposely obscure fashion, as if they wanted to keep the plebs away from their holy secrets. This document instead begs to be read; some parts are more in the style of a novel than of a technical document.
- bsder 2y agoAI has yet to demonstrate that it can do anything different from what a group of people could sit down and do. Sure, the AI may be able to do it faster, but there hasn't yet been anything demonstrated that exceeds what humans can do. If it would be illegal for a group of people to do something, it is also going to be illegal for an AI do so. Why is that so surprising?
- musicale 2y ago> "we trained on copyrighted documents, and made a general purpose AI, and then people paid to use our AI to compete with the people who owned the documents" This is a good distillation. A bit like "we trained our system on various works of art and music, and now it is being sold as a service that competes with the original artists and musicians."