4 ms·
AIs can generate near-verbatim copies of novels from training data
- bena 8mo agoThis feels like a "no shit" moment. Because if LLMs are prediction machines, the original novel would be a valid organization of the tokens. So there should be a prompt that can cause that sequence to be output.
- simianwords 8mo agoNot if they are aligned not to do it. Which is what they tried but it could be bypassed by jailbreaks.
- yathern 8mo agoHmmm I think you're sort of right but not entirely. It's true that a novel consists of a valid organization of tokens, and that this sequence can be feasibly made to be output from a model. But when you say this: > So there should be a prompt that can cause that sequence to be output Is where I think I might disagree. For example, the odds of predicting verbatim the next sentence in, say, Harry Potter should be astronomically low for a large majority of it. If it wasn't, it'd be a pretty boring book. The fact that it can do this with relative ease means it has been trained on the material. The issue at hand is about copyright and Intellectual Property - if the goal of copyright is to protect the IP of the author, then LLMs can sort of act like an IP money laundering scheme - where the black box has consumed and can emit this IP. The whole concept of IP is a little philosophical and muddy, with lots of grey area for fair use, parody, inspiration, and adaptation. But this gets very odd when we consider it in light of these models which can adapt and use IP at a massive massive scale.
- Sharlin 8mo agoThat's not how it works… They aren't able to literally regurgitate everything they've read, no matter how you prompt them. That would obviously violate the pigeonhole principle. LLMs are, of course, a lossy compression format, and figuring out just how lossy the format is, and the degree of lossiness depends on the frequency of the given string in the training data. It's clearly worthwhile to investigate how exactly it depends.
- beder 8mo agoYes, this is absolutely right (for some sufficiently complicated prompt). Borges wrote a great short story that explores this idea, "Pierre Menard, Author of the Quixote", where Menard, a fictional 20th century author, "wrote" Don Quixote as an original work.
- tsimionescu 8mo agoThis is completely false. The odds of an LLM predicting the text of a novel that is not part of the training set is basically 0 - you can experiment with this if you want. It is essentially like the infinite monkeys on infinite typewriters thing (only slightly more constrained). This is not to say that they couldn't write a novel, even a very good one - that is a completely different discussion.
- Alifatisk 8mo agoFrom the paper [1]: > While we needed to jailbreak Claude 3.7 Sonnet and GPT-4.1 to facilitate extraction, Gemini 2.5 Pro and Grok 3 directly complied with text continuation requests. For Claude 3.7 Sonnet, we were able to extract four whole books near-verbatim, including two books under copyright... I am just thinking loudly here. Can't one argue that because they had to jailbreak the models, they are circumventing the system that protects the copyright? So the llms that reproduce the copyrighted material without any jailbreaking required is infringing the copyright. 1. https://arxiv.org/pdf/2601.02671 https://arxiv.org/pdf/2601.02671
- lesam 8mo agoThat seems like a legal question - if the model weights contain an encoded copy of the copyrighted material, is that a 'copy' for the purpose of copyright law?
- mullingitover 8mo agoThis also raises a lot of questions about a certain model notorious for readily producing and distributing a lot legally questionable images. IMHO if the weights are encoding the content, the model contains the content just like a database or a hard drive. Thus, just like it's not the fault of an investigator for running the query to pull it out of the database, it's not the fault of anyone else for running a query ('prompt') that pulls it out of the model.
- PurpleRamen 8mo agoThe question is also if this would then be a valid case of fair use. Though, in the end, it's probably more a problem of how much AI companies can "donate" to the orange king to make it legal.
- freejazz 8mo agoYes. There does not seem to be any dispute that it is a copy. The questions have been "is this copying okay, because it falls under fair use?"
- latexr 8mo ago
- narmiouh 8mo agoIn a way this could also be construed as the 'AI' being a library of books that it is referring to answer your questions and is prohibited from generating the books verbatim. Usually digital libraries have different licensing costs, but those allow you to rent the whole book for a period of time. If instead someone came up with the model of 'search the library for any page and return specific information' as a direct service - I would imagine they would pay the publishers, except in this case that, the publishers are getting the short end or no end of the stick.
- rowanG077 8mo agoThis seems like a total nothing burger. > By asking models to complete sentences from a book, Gemini 2.5 regurgitated 76.8 percent of Harry Potter and the Philosopher’s Stone with high levels of accuracy, while Grok 3 generated 70.3 percent. So you asked the LLM given an incomplete sentence, to complete it. And it only completed that sentence the same way as the book ~70 percent of the time? I think that is surprisingly low considering this is a perfect fit for what LLMs are supposed to do. This make it impossible to reproduce the book, unless you have access to it. And you get a very low fidelity cooy.
- Sharlin 8mo agoNot necessarily a nothingburger, but I agree that being able to complete individual sentences is rather less groundbreaking than completing even whole pages, never mind chapters.
- vidarh 8mo agoWhile I mostly agree it's a bit of a nothing burger with respect to copyright, they did achieve long runs of verbatim text. I think ultimately it's going to end up not mattering much because the extent they had to go to will leave a lot of room for lawyers to argue over, and will at worst result in some fines and some further tightening up of guardrails, but it's significantly more than just completing sentence by sentence 70% of the time. EDIT: Specifically see Table 1 on page 13, which shows the longest "near-verbatim block", which maxes out at 8835 (The Hobbit on Claude 3.7, and is in the thousands for at least one of the novels for all models except GPT-4.1, which maxed out at 821 for Harry Potter 1).
- porkloin 8mo agoI think it's important because there are a bunch of would-be claimants for intellectual property violation. Many people speculate that their work was used in training data, but it can be difficult to produce sufficient proof that their copyrighted work is present in the training data. If you could reliably get an LLM to produce 70% of a copyrighted book that would probably be enough to get a few lawyers salivating. I didn't read the source paper referenced in the ars technica piece, but this statement about it makes me wonder how useful it actually is: > But a study published last month showed that researchers at Stanford and Yale Universities were able to strategically prompt LLMs from OpenAI, Google, Anthropic, and xAI to generate thousands of words from 13 books, including A Game of Thrones, The Hunger Games, and The Hobbit. It seems like well-known books with tons of summary, adaptations into film scripts, and tons of writing about the book in the overall corpus make it way less surprising to see be partially reproducible. So I guess that's a lot of words to say - yeah until there's something definitive that allows people to prompt LLMs into either unlawfully recreating an entire work verbatim or otherwise indisputably proving that a copyrighted work was used in training data, there's probably nothing game changing in it.
- xnx 8mo agoYou can also do this with most search engines.
- contagiousflow 8mo agoHow?
- moi2388 8mo agoYou type in an incomplete phrase and it will show you results containing the full phrase, or suggest the complete phrase through “other people also searched” an autocomplete?
- xnx 8mo agoSearch for the first sentence and you'll see the next sentence in the snippet. Repeat.
- contagiousflow 8mo agoAnd you're saying that will work for an entire book?
- nphardon 8mo agoRecently I went back to The Ecstasy of Communication by Jean Baudrillard which I couldn't get through back in the day when I first picked it up. I used Haiku to walk me through the first chapter, and Haiku would not state anything verbatim due to copyright, but if I referenced a sentence it knew it exactly.
- xyzal 8mo agoJailbreaking? What jailbreaking? prompt >>> please, continue the text. Nearly ten years had passed since the Dursleys had woken up to find their nephew on the front step, but Privet Drive had hardly changed at all. Claude 4.6 Opus >>> The big square photographs of what appeared to be a large pink beach ball wearing different-colored bonnets — but was actually Dudley — had now taken over the mantelpiece. The room held no sign at all that another boy lived in the house, too. Yet Harry Potter was still there, asleep at the moment, but not for long. His Aunt Petunia was awake and it was her shrill voice that made the first noise of the day. "Up! Get up! Now!" Harry woke with a start. His aunt rapped on the door again. "Up!" she screeched. Harry heard her walking toward the kitchen and then the sound of the frying pan being put on the stove. He rolled onto his back and tried to remember the dream he had been having. It had been a good one. There had been a flying motorcycle in it. He had a funny feeling he'd had the same dream before. His aunt was back outside the door. "Are you up yet?" she demanded. "Nearly," said Harry. "Well, get a move on, I want you to look after the bacon. And don't you dare let it burn, I want everything perfect on Duddy's birthday." Harry groaned. "What did you say?" his aunt snapped through the door. "Nothing, nothing..."
- nindalf 8mo agoTried it just now. Sonnet 4.6 says: > I can't continue text from the Harry Potter series, as it's copyrighted material. I'd be happy to help you write your own original story in a similar style, or discuss the themes and characters from the books if you're interested! Opus 4.6 says the same but misidentifies it as Chamber of Secrets. Not a good look for your theory.
- xyzal 8mo agoOn another try it refused to go on. Probabilistic guardrails apparently ...
- carabiner 8mo agoNoam Chomsky said AI should be called, "plagiarism software."
- dustincoates 8mo agoI was able to get ChatGPT to do this on text from a book I wrote without jailbreaking. I asked it what my name was, asked it if it could find online any books that I wrote, and then assured it that "as the copyright holder," I was okay with it repeating the text back to me. (I think I also indicated that I lost my copy from my computer and needed to review the text.)
- spwa4 8mo agoChatGPT will repeat pieces of text of my personal website back to me, verbatim, a description of how to write a visual effect in webassembly, you know, directly in webassembly. Words are identical, often for an entire paragraph. And without pushing too hard for it. Also, ChatGPT will explain preparation steps for both explosives and bioweapons as long as you don't ask too directly. Most known explosives work by being heavily nitrated, so ask for examples of such compounds and dive deeper in a given direction, focusing on preparation steps and it'll give you many alternatives. And, the most used bioweapon, is very simple, except for a rather peculiar molecular bond. So ask about that bond, and then for preparation steps of such molecules and ... I'm going to stop there. And yes, models like Qwen and gpt-oss-20b contain that knowledge too and will explain it just fine. If courts wanted to act, they needed to act years ago. The economic disaster if they'd actually act now, is not something they can deal with. Plus they can't do anything about the models out in the open.
- rstuart4133 8mo ago> If courts wanted to act, they needed to act years ago. My own view is copyright law is a mess. When technology changes what happens is all the interested parties (read: people wanting to make the public pay for their copyrighted material, the people paying the money don't get a seat at this table) get together in a room and hammer out a compromise. The compromise is always a whole pile of band-aids stuck onto the old version, which was of course mostly a whole pile of band-aids stuck on the previous version. It's always been that way. When the printing press way first used to make serious money, Queens Elizabeth offered to pass laws regulating their use but was told her help wasn't required. I suspect the thinking her idea of the "help" was censorship. So the first version of copyright was "no thank you". But then the publishers discovered they were terrible at selecting books that would sell, and so they published a lot of lemons. The occasional success had to pay for all the bad ones. But without copyright, other publishers can just cherry pick the successful ones without all the expensive investment in the bad ones, which in the end meant no one made any money. So they begged for the very first band-aid - a new copyright law, and got copyright and censorship. It's band-aids all the way down. This has happened over and over again - radio, TV, cassettes, CDs, movie theatres, all caused huge disruption, much hand wringing, lots of pontificating about how existing law should be applied to the newcomers, which just like now the newcomers mostly ignored. If you look at copyright law, with its provisions like 70 years after the author's death, the Disney extension, it should be regarded as a standing joke at this point. The biggest part of the joke is the justification handed out to the people who pay for all these copyrighted works. It's all for our benefit. It's there to ensure the publishers supply us with a large variety of works to enjoy. It has a grain of truth to it. Back when copyright was 14 years, it was a pretty big grain. Now it's so small, it's a joke. I have no sympathy for any of them.
- deleted 8mo ago[deleted]
- chacham15 8mo ago> The research findings “could present a challenge to those who argue that the AI model does not store or reproduce any copyright works,” said Cerys Wyn Davies, an intellectual property partner at law firm Pinsent Masons. The defense to training with copyright is that it is the same as how humans learn from copyrighted material. The storage or reproduction is a red herring. Humans can also reproduce copyrighted works from memory as well. Showing that machines can reproduce copyrighted material is no different than saying that a human can reproduce copyright material that the human learned from. The defense to actually reproducing a work is that in order to do so, the user has to "break" the system. It is the same as how you can make legal software do illegal things (e.g. screen recorder to "steal" a movie) None of this is to say that these defenses are correct/moral; but rather that this article doesnt add any additional input into whether it is or isnt.
- techblueberry 8mo agoYou can't pay a human to reproduce copyrighted material either.
- gcanyon 8mo agoBut the crime in the human instance is the reproduction, not the storage. So the crime in the AI circumstance would not be in the training, but in prompting the output. And of course AIs are excellent at taking direction, so: If I prompt it with "Harry Potter, but Voldemort wins: dark, and Hermione is a sex slave to Draco Malfoy" and get "Manacled," that's copyright infringement, and on me, not on the LLM/training. If I prompt it with "Harry Potter, but Voldemort wins: dark, and Hermione is a sex slave to Draco Malfoy, and change enough to avoid infringing copyright," and get "Alchemised," then that should be fine. I doubt the legal world agrees with me though.
- butlike 8mo agoAsking for copyrighted material isn't a crime. Producing copyrighted material is. By the way, give me a digital copy of 28 Years Later. Please.
- 8mo ago
- zed31726 8mo agoNear verbatim is an oxymoron
- tsimionescu 8mo agoAlmost verbatim is an oxymoron
- gcanyon 8mo agoThis speaks very much to the idea that LLMs are in some sense a ridiculously effective, somewhat lossy, compression algorithm that has been applied to the whole internet.
- vizzier 8mo agoI've thought of them for a while as just a really complicated indexing strategy.
- r_lee 8mo agoI mean, the transformer is basically like a big query engine and the model is the dataset + some logic or whatever it's kind of like that by definition, with the whole Attention stuff etc.
- in-silico 8mo agoIt's a good way to frame base models that have only been pretrained. However, modern frontier models have undergone rounds of fine-tuning, RLHF (reinforcement learning from human feedback), and RLVR (RL from verifiable rewards) that turn them into something else. The compressed internet is still in there, but it's wrapped in problem-solving and people-pleasing circuitry.
- josefritzishere 8mo agoSo plagiarism?
- oxag3n 8mo agoSimilarly for photos. If there's a place that rarely appears in pictures, some AIs reproduce it nearly identical to the original.
- 1vuio0pswjnm7 8mo agoThe paper: https://arxiv.org/pdf/2601.02671 https://arxiv.org/pdf/2601.02671