12 ms·
Thomson Reuters wins first major AI copyright case in the US
- EnderWT 2y agohttps://archive.is/mu49I https://archive.is/mu49I
- NewsaHackO 2y agoHow does this affect LLM systems that already have their corpus integrated?
- trod1234 2y agoThe judge ruled this as a violation of copyright. Its the same as hosting any copyright material absent a valid license, criminal copyright piracy. They would need to figure out a way to prune the respective weights so that such material is not available, or risk legal fury.
- Workaccount2 2y agoThey would just censor output. Youtube doesn't need to figure out how to stop copyright material from being uploaded, they need to stop it from being shared.
- anticensor 2y ago> They would need to figure out a way to prune the respective weights so that such material is not available, or risk legal fury. You want to reliably train it away from outputting the undesired outputs, not keeping it ignorant about them.
- preinheimer 2y agoGreat. The stated goal of a lot of these companies seems to be “train the model on the output of humans, then hire us instead of the humans”. It’s been interesting that media where watermarking has been feasible (like photography) have seen creators get access to some compensation, while text based creators get nothing.
- rolph 2y agorotate similar [but different] fonts [or character pages] over each character. the sequence represents data thus watermark.
- WillAdams 2y agobut the font changes won't be expressed in the (plain text) output of the LLM.
- yifanl 2y agoPresumably the font will represent letters to look like a different letter, making it not useful to LLMs scraping the site but useful for visual readers. This would have detrimental effects to people who use screen readers or have their own stylesheets of course.
- WillAdams 2y agoFor that it would make more sense to run a routine which replaces letters with visually identical glyphs at different encoding points.
- palmotea 2y ago> For that it would make more sense to run a routine which replaces letters with visually identical glyphs at different encoding points. It seems like that would be pretty easily defeatable with the similar mapping to the one used to do the replacement.
- WillAdams 2y agoYes, but it's one more hurdle for them.
- chefandy 2y agoThat would only stymie the smallest-time players. Things like sideways text in margins or rotated table column headers are common enough that these have been solved problems for decades. Breaking the text down into specific elements and handling it differently or ignoring it altogether based on context and content is trivial.
- oidar 2y agoRoss intelligence was creating a product that would directly compete against Thomson Reuters. Pretty clearly not fair use.
- YesBox 2y agoThanks. The article wasn't loading for me, just the headline and image and footer. I was about to leave thinking that's all there is.
- dang 2y agoWe detached this comment from https://news.ycombinator.com/item?id=43018356 https://news.ycombinator.com/item?id=43018356. Your post is totally fine; I just want to save space at the top of the thread (where the parent is now pinned).
- varsketiz 2y agoGreat decision for humans. Is this type of risk the reason why OpenAI masquerades as a non-profit?
- deleted 2y ago[deleted]
- veggieroll 2y ago> Thomson Reuters prevailed on two of the four factors, but Bibas described the fourth as the most important, and ruled that Ross “meant to compete with Westlaw by developing a market substitute.” Yep. That's what people have been saying all along. If the intent is to substitute the original, then copying is not fair use. But the problem is that the current method for training requires this volume of data. So the models are legitimately not viable without massive copyright infringement. It'll be interesting to see how a defendant with a larger wallet will fare. But this doesn't look good. Though big-picture, it seems to me that the money-ed interests will ensure that even if the current legal landscape doesn't allow LLM's to exist, then they will lobby HARD until it is allowed. This is inevitable now that it's at least partially framed in national security terms. But I'd hope that this means there is a chance that if models have to train on all of human content, the weights will be available for free to all humans. If it requires massive copyright infringement on our content, we should all have an ownership stake in the resulting models.
- saulpw 2y ago> So the models are legitimately not viable without massive copyright infringement. Copyright is not about acquisition, it is about publication and/or distribution. If I get a copy of Harry Potter from a dumpster, I can read it. If a company gets a copy of *all books from a torrent, they can use it to train their AI. The torrent providers may be in violation of copyright, and if the AI can be used to reproduce substantive portions of the original text, the AI companies then may be in violation of copyright, but simply training a model on illegally distributed text should not be copyright infringement.
- veggieroll 2y agoI mean, you're right in the abstract. If you train an LLM in a void and never do anything with the model, sure. But that's not what anyone is doing. People train models so that someone can actually use them. So I'm not sure how your comment is helpful other than to point out that distinction (which doesn't make much difference in this case specifically or how copyright applies for LLM's in general)
- deleted 2y ago
- iandanforth 2y agoEstablishing precedent by defeating an already dead company in court is neither impressive nor likely to hold up for other companies.
- asadotzler 2y agoincorrect
- teruakohatu 2y agoRoss Intelligence was more a search interface with natural language and, probably, vector based similarity. So I suspect they were hosting and using the corpus in production, not just training a model on it.
- dkjaudyeqooe 2y agoThe fair use aspect of the ruling should send a chill down the spines of all generative AI vendors. It's just one ruling but it's still bad.
- palmotea 2y ago> The fair use aspect of the ruling should send a chill down the spines of all generative AI vendors. It's just one ruling but it's still bad. So, in other words, it's good.
- 2OEH8eoCRo0 2y agoFantastic news!
- JackC 2y agoHere's the full decision, which (like most decisions!) is largely written to be legible to non-lawyers: https://storage.courtlistener.com/recap/gov.uscourts.ded.72109/gov.uscourts.ded.72109.770.0.pdf https://storage.courtlistener.com/recap/gov.uscourts.ded.721... The core story seems to be: Westlaw writes and owns headnotes that help lawyers find legal cases about a particular topic. Ross paid people to translate those headnotes into new text, trained an AI on the translations, and used those to make a model that helps lawyers find legal cases about a particular topic. In that specific instance the court says this plan isn't fair use. If it was fair use, one could presumably just pay people to translate headnotes directly and make a Westlaw competitor, since translating headnotes is cheaper than writing new ones. And conversely if it isn't fair use where's the harm (the court notes no copyright violation was necessary for interoperability for example) -- one can still pay people to write fresh headnotes from caselaw and create the same training set. The court emphasizes "Because the AI landscape is changing rapidly, I note for readers that only non-generative AI is before me today." But I'm not sure "generative" is that meaningful a distinction here. You can definitely see how AI companies will be hustling to distinguish this from "we trained on copyrighted documents, and made a general purpose AI, and then people paid to use our AI to compete with the people who owned the documents." It's not quite the same, the connection is less direct, but it's not totally different.
- echelon 2y agoIf the copyright holders win, the model giants will just license. This effectively kills open source, which can't afford to license and won't be able to sublicense training data. This is very bad for democratized access to and development of AI. The giants will probably want this. The giants were already purchasing legacy media content enterprises (Amazon and MGM, etc.), so this will probably further consolidation and create extreme barriers to entry. If I were OpenAI, I'd probably be very happy right now. If I were a recent batch YC AI company, I'd be mortified.
- dkjaudyeqooe 2y agoLicense what? Every available copyrighted work? Even getting a tiny fraction is not practical. To the contrary, this just means companies can't make money from these models. Those using models for research and personal use wouldn't be infringing under the fair use tests.
- Animats 2y agoThis isn't really about "AI". It's about copying summaries. Google was fined for this in France for copying news headlines into their search results, and now has to pay royalties in the EU. Westlaw is a summarizing and indexing service for court case results. It's been publishing that info in book form since 1872. Ross was trying to compete with Westlaw, but used Westlaw as an input. West's "Key Numbers" are, after a century and a half, a de-facto standard.[2] So Ross had to match that proprietary indexing system to compete. Their output had to match Westlaw's rather closely. That's the underlying problem. The court ruled that the objective was to directly compete with Westlaw, and using Westlaw's output to do that was intentional copyright infringement. This looks like a narrow holding, not one that generally covers feeding content into AI training systems. [1] https://apnews.com/article/google-france-news-publishers-copyright-7a7e484f55297e9803d17f736ff923a0 https://apnews.com/article/google-france-news-publishers-cop... [2] https://guides.law.stanford.edu/cases/keynumbersystem https://guides.law.stanford.edu/cases/keynumbersystem
- mmooss 2y agoTR may have intentionally chosen an easy battle to begin their legal war.
- ascorbic 2y agoThey began this case in 2020, before any of the most important models existed
- zozbot234 2y agoThe case involves headnotes, not just key numbers. Your links provide examples of such headnotes, which make it very clear that a lot of human creativity and judgment is involved in authoring them - they're not a matter of purely factual information, such as a phonebook. Thus, the headnotes are copywritten, and translating them to a different language doesn't negate that copyright. This looks like a slam dunk case, but it has very little to do with AI training as such - the AI was only used to create a kind of rough indexing over the translated text. If this was only about key numbers, it might have gone the other way because the fact-like element there is considerably greater.
- MonkeyClub 2y agoFrom p. 6: "But a headnote can introduce creativity by distilling, synthesizing, or explaining part of an opinion, and thus be copyrightable." Does this set a precedent, whereby AI-generated summaries are copyrightable by the LLM owners?
- rvz 2y agoSee. The fair-use excuses that the AI proponents here were trying to hang on to for dear life have fallen flat on this ruling. This is going to be one of many cases in which there will be licensing deals being made out of this to stop AI grifters claiming 'fair use' to try to side-step copyright laws because they are using a gen AI system. OpenAI ended up paying up for the data with Shutterstock and other news sources. This will be no different.
- Salgat 2y agoMy biggest concern is, what happens when countries like China, who aren't restricted by this, far outpace western countries in this technology? Do we just shrug and accept our far inferior models? LLMs are a productivity multiplier (similar to a search engine), so it'll have a large impact on the economy if licensing costs prohibit large scale training.
- esafak 2y ago"Our" models are not inferior. There is plenty of data, and the next frontier is prediction-time compute and data synthesis. Shouldn't the Chinese worry that they are depressing the publication of commercial IP?
- Salgat 2y agoOur models are not inferior because we're not constrained by copyrights, which is the entire concern we're discussing.
- blibble 2y agoonce the case law is set, I look forward to suing everyone that's ever trained a model for $300,000 PER WORK each time they ingested my code from GitHub whoever wrote those indemnity policies is going to regret it
- anticensor 2y ago> once the case law is set, I look forward to suing everyone that's ever trained a model for $300,000 PER WORK each time they ingested my code from GitHub Didn't you already share it on GitHub royalty-free?
- mmooss 2y agoThomson Reuters chose to sue Ross Intelligence, not a company like Google or even OpenAI. I wonder how deeper pockets would have affected the outcome. I wonder how the politics played out. The big AI companies could have funded Ross Intelligence, who could have threatened to sabotage their legal strategies by tanking and settling their own case in TR's favor.
- r00fus 2y agoWhat's to say that's not the next step? First step, stop your competitors who are copying your IP. Lawyers are gonna be happy is my thought.
- mmooss 2y agoI missed this line from the article: Even before this ruling, Ross Intelligence had already felt the impact of the court battle: the startup shut down in 2021, citing the cost of litigation.
- rudedogg 2y agoIt being in the legal realm probably had some impact. These tools can be seen as an attack on your profession, and I’m sure that affects the decision, whether conscious or not.
- mmooss 2y agoThe judge is a user of Westlaw and not a shareholder, as is the judge's office. They would like Westlaw to be cheaper and easier to use. Westlaw takes the judge's work product and profits from it. Arguably, Westlaw should demand all judges recuse themselves!
- ars 2y agoIt would be quite an interesting result if we could have true General AI, but we don't simply because of copyright. I'm aware this isn't a concern yet, but imagine if the future played out this way.... Or worse: Only those with really deep pockets can pay to get AI, and no one else can, simply because they can't afford the copyright fees.
- simonw 2y agoInteresting to note from this 2020 story (when ROSS shut down) that the company was founded in 2014 and went out of business in 2020: https://www.lawnext.com/2020/12/legal-research-company-ross-to-shut-down-under-pressure-of-thomson-reuters-lawsuit.html https://www.lawnext.com/2020/12/legal-research-company-ross-... The fact that it took until 2024 for the case to resolve shows how long the wheels of justice can take to turn!
- qingcharles 2y agoLitigation takes forever. Especially when you factor COVID in. I'm still litigating cases from over a decade ago that are probably several years from resolution, just in the district court. Then you can spin through appeals courts for another five years. And that's civil. Criminal, especially a death row case, can take 20+ years to exhaust every level of appellate review. In Illinois there are at least nine levels of review available to you without going through second rounds of review, state habeas, and collateral attacks like applications for clemency, pardons etc. If you're not paying for lawyers, expect each level to take around two years or more.
- freeAgent 2y agoMy father practiced corporate tax law and regularly had cases at trial that resolved issues from 20-30 years prior.
- simonw 2y agoThat's wild, I had no idea. I have trouble imagining a case where it's worth spending 30 years coming to a conclusion, but I guess that's one many reasons I'm not a corporate tax lawyer!
- freeAgent 2y agoOne such repeating case ended up settling for over $10B, so it was definitely worth it! To clarify, they spent decades litigating the same fundamental issue for each year’s tax filings, with each filing year taking multiple years to get to court. The plaintiffs won every single case until the government finally settled all the remaining tax years for that amount. Each year prior was worth hundreds of millions.
- jug 2y agoI spontaneously feel like this is bad news for open AI, while playing in the hands of corporate behemoths able to strike expensive deals with major publishers and top it off with the public domain. I’m not sure this signals the end of AI and a victory for the human, but rather who gets to train the models?
- lazycog512 2y agoseems like delaware can't scare tech companies out of re-incorporating any faster
- aurizon 2y agoAt the heart of this is a very greedy racket:- court reporters who 'own' the copyright to every word spoken by anyone in court that they transcribe to a transcript that they do not own the source to (judges/witnesses/lawyers/defendants in truth own it) They then milk huge fees for these transcripts and limit use/access/derivative works with huge fees. An AI verbatim transcriber would up end them, so that will be prevented, as will anything that shakes the tree.
- habinero 2y agoNo, their work is valuable and they deserve to make money off of it. The reason why it's valuable is it's transcribed live (usually with video) and is accurate and verifiable. Words and names are spelled correctly and speakers are correctly identified. Court reporters will stop speakers and ask for spelling or to repeat words. AI transcriptions can't do that.
- aurizon 2y agoIt is valuable, but it should not be extortionable
- vaadu 2y agoDoes anyone think Deepseek or other non-western AIs will respect copyright? This is going to make Deepseek and its kin much more valuable.
- jll29 2y agoNote this case is explicitly NOT about large language model type AI - Ross' product is just a traditional search engine (information retrieval system), not a neural transformer a la ChatGPT. About judge Bibas: https://en.wikipedia.org/wiki/Stephanos_Bibas https://en.wikipedia.org/wiki/Stephanos_Bibas
- nickpsecurity 2y agoAlmost every article I read on fair use talked like I could only use small amounts while not competing with them. AI people focus on a tiny number of precedents that they stretch very far. A reasonable person wouldn’t come up with their interpretation of fair use after looking at how most examples play out in court. It shouldn’t surprise the writer that the AI companies’ versions of fair use didn’t hold much weight. They should assume that would be true. Then, be surprised any time a pro-AI ruling goes against common examples in case law. The AI companies are hoping to achieve that by throwing enough money at the legal system.
- afarviral 2y agoIf those 4 aspects are used to judge whether "fair use", I'd say that's the nail in the coffin, because of course it isn't fair use and that's totally fair. Here I was thinking "transformative" was somehow a sticking point in all this.
- linkerdoo 2y ago[dead]
- xyzal 2y agoI can't understand how some commenters frame such a result as not good. The big players will have no problem licensing large corpora to train their models, while my tiny site won't be vacuumed (legally at least) by scrapers if I won't agree. My willingness to upload my projects anywhere is in the historical lows given the current state, honestly.
- gradientsrneat 2y agoWestlaw is to the legal profession what ResearchGate and others are to science research. They profit from information from the commons, and charge as much as the market will bear. Only one of the many reasons the legal profession is so expensive.
- biohcacker84 2y agoIf copyright forces a diversity of AIs. That would be good. Every AI company using its own created training, resulting in AIs that are similar but not identical, is in my opinion much better than one or very few AIs.
- KimReyes 2y ago[dead]
- KimReyes 2y ago[dead]