6 ms·
Simon Willison had an analysis of Claude's system prompt back in May. One of the things that stood out was the effort they put in to avoiding copyright infring
by loudmax 11mo ago
Simon Willison had an analysis of Claude's system prompt back in May. One of the things that stood out was the effort they put in to avoiding copyright infringement: https://simonwillison.net/2025/May/25/claude-4-system-prompt/#seriously-don-t-regurgitate-copyrighted-content https://simonwillison.net/2025/May/25/claude-4-system-prompt...
Everyone knows that these LLMs were trained on copyrighted material, and as a next-token prediction model, LLMs are strongly inclined to reproduce text they were trained on.
- miltonlost 11mo agoAll AI companies know they're breaking the law. They all have prompts effectively saying "Don't show that we broke the law!". That we continue to have tech companies consistently breaking the law and nothing happens is an indictment of our current economy.
- mock-possum 11mo agoI don’t read this as “don’t show we broke the law,” I read it as “don’t give the user the false impression that there’s any legal issue with this generated content.” There’s nothing law breaking about quoting publicly available information. Google isn’t breaking the law when it displays previews of indexed content returned by the search algorithm, and that’s clearly the approach being taken here.
- Q6T46nT668w6i3m 11mo agoMasked token prediction is reconstruction. It goes far beyond “quoting.”
- lokar 11mo agoThe whole industry is based on breaking the law. You don’t get to be Microsoft, Google, Amazon, meta, etc without large amounts of illegality. And the VC ecosystem and valuations are built around this assumption.
- admaiora 11mo agoAnd it's a question of do we accept breaking law for the possibility to have the greatest technological advancement of the 21st century. In my opinion, legal system has become a blocker for a lot of innovation, not only in AI but elsewhere as well.
- saghm 11mo agoWithout agreeing or disagreeing with your view, I feel like the the issue the issue with that paradigm is inconsistency. If an individual "pirates", they get fines and possible jail time, but if a large enough company does it, they get rewarded by stockholders and at most a slap on the wrist by regulators. If as a society we've decided that the restrictions aren't beneficial, they should be lifted for everyone, not just ignored when convenient for large corporations. As it stands right now, the punishments are scaled inversely to the amount of damage that the one breaking the law actually is capable of doing.
- rpdillon 11mo agoThis is a point that I don't see discussed enough. I think anthropic decided to purchase books in bulk, tear them apart to scan them, and then destroy those copies. And that's the only source of copyrighted material I've ever heard of that is actually legal to use for training LLMs. Most LLMs were trained on vast troves of pirated copyrighted material. Folks point this out, but they don't ever talk about what the alternative was. The content industries, like music, movies, and books, have done nothing to research or make their works available for analysis and innovation, and have in fact fought industries that seek to do so tooth and nail. Further, they use the narrative that people that pirate works are stealing from the artists, where the vast majority of money that a customer pays for a piece of copyrighted content goes to the publishing industry. This is essentially the definition of rent seeking. Those industries essentially tried to stop innovation entirely, and they tried to use the law to do that (and still do). So, other companies innovated over the copyright holder's objections, and now we have to sort it out in the courts.
- Q6T46nT668w6i3m 11mo agoI don’t follow. You’re punishing the publishing industry by punishing authors?
- blibble 11mo agoand training on mountains of open source code with no attribution is exactly the same the code models should also be banned, and all output they've generated subject to copyright infringement lawsuits the sloppers (OpenAI, etc) may get away with it in the US, but the developed world has far more stringent copyright laws and the countries that have massive industries based on copyright aren't about to let them evaporate for the benefit of a handful of US tech-bros
- terminalshort 11mo agoNo thank you. I am perfectly fine with AI training on my open source code and it is perfectly legal because my open source code does not include a license that bans AI training.
- blibble 11mo agowhich license is that then? because other than public domain they all require at least displaying the license, which "AI" ignores
- Workaccount2 11mo agoTraining on copyright is not illegal. Even in the lawsuit against anthropic it was found to be fair use. Pirating material is a violation of copyright, which some labs have done, but that has nothing to do with training AI and everything to do with piracy.
- dahart 11mo agoThere is US precedent for training being deemed not fair use. https://www.dglaw.com/court-rules-ai-training-on-copyrighted-works-is-not-fair-use-what-it-means-for-generative-ai/ https://www.dglaw.com/court-rules-ai-training-on-copyrighted... Why wouldn’t training be illegal? It’s illegal for me to acquire and watch movies or listen to songs without paying for them*. If consuming copyrighted material isn’t fair use, then it doesn’t make sense that AI training would be fair use. * I hope it’s obvious but I feel compelled to qualify that, of course, I’m talking about downloading (for example torrenting) media, and not about borrowing from the library or being gifted a DVD, CD, book or whatever, and not listening/watching one time with friends. People have been successfully prosecuted for consuming copyrighted material, and that’s what I’m referring to.
- terminalshort 11mo agoThat interpretation is not correct. The owner explicitly denied license to the data and then the company went to a third party to gain access to the data that they were denied license to. > When building its tool, Ross sought to license Westlaw’s content as training data for its AI search engine. As the two are competitors, Thomson Reuters refused. Instead, Ross hired a third party, LegalEase, to provide training data in the form of “Bulk Memos,” which were created using Westlaw headnotes. Thomson Reuters’s suit followed, alleging that Ross had infringed upon its copyrighted Westlaw headnotes by using them to train the AI tool.
- dahart 11mo agoYou’re contradicting the conclusion / interpretation written on dglaw.com? What is incorrect, exactly? It doesn’t seem like your summary challenges either my comment or the article I linked to, it’s not clear what you’re arguing. The court did find in this case that the use of the unlicensed data used for AI training was not fair use.
- terminalshort 11mo agoThis is incorrect. Two judges have now ruled that training on copyrighted data is fair use. https://www.whitecase.com/insight-alert/two-california-district-judges-rule-using-books-train-ai-fair-use https://www.whitecase.com/insight-alert/two-california-distr...
- hulitu 11mo agoYou can always vote, but there is always someone going through the back door paying politicians and judges.
- qustrolabe 11mo agopost trained models strongly inclined to pass response similar to what got them high RL score, it's slightly wrong to keep thinking of LLMs as just next token predictions from dataset's probability distribution like it's some Markov Chain