12 ms·
A federal judge sides with Anthropic in lawsuit over training AI on books
- 3PS 1y agoBroadly summarizing. This is OK and fair use: Training LLMs on copyrighted work, since it's transformative. This is not OK and not fair use: pirating data, or creating a big repository of pirated data that isn't necessarily for AI training. Overall seems like a pretty reasonable ruling?
- simmerup 1y agoDepends whether you actually agree its transformative
- lesuorac 1y agoFor textual purposes it seems fairly transformative. If you train a LLM on harry potter and ask it to generate a story that isn't harry potter then it's not a replacement. However, if you train a model on stock imagery and use it to generate stock imagery then I think you'll run into an issue from the Warhol case.
- johnnyanmac 1y agoThe nature of how they store data makes it not okay in my books. You massage the data enough and you can generate something that seems infringement worthy.
- ticulatedspline 1y agoFor closed models the storage problem isn't really a problem, they can be judged by what they produce not how they store it as you don't have access to the actual data. That said, open weight LLMs are probably screwed, if enough of the work remains in the weights such that they can be extracted (even if it's without even talking to the LLM) then the weight file itself represents a copy of the work that's being distributed. So enjoy these competent run-at-home models while you can, they're on track for extinction.
- ninetyninenine 1y agoWhy doesn’t this apply to humans? If I memorize something such that it can be extracted did I violate the law? It’s only if I choose to allow such extraction to occur then I’m in violation of the law right? So if I or an LLM simply doesn’t allow said extraction to occur, memorization and copying is not against the law.
- ranger_danger 1y agoI think an important distinction here is distribution... did you tell someone else what you memorized? Is downloading a model akin to distributing that same information?
- ninetyninenine 1y agoWhat if I don't download the model and I just communicate with it. Sort of like chatting with another human. That's not a copyright issue right? I mean that's how most LLMs are deployed today.
- ranger_danger 1y agoMy understanding is that it depends on a judge/jury's subjective opinion on how similar the output is to something copyrightable. Perhaps intent may play a role as well.
- deleted 1y ago[deleted]
- deleted 1y ago[deleted]
- ranger_danger 1y agoI wonder if https://en.wikipedia.org/wiki/Illegal_number https://en.wikipedia.org/wiki/Illegal_number comes into play here.
- sidewndr46 1y agoWasn't that just over an arrangement of someone else's photographs?
- lesuorac 1y agohttps://en.wikipedia.org/wiki/Andy_Warhol_Foundation_for_the_Visual_Arts,_Inc._v._Goldsmith https://en.wikipedia.org/wiki/Andy_Warhol_Foundation_for_the... I wouldn't call it that. Goldsmith took a photograph of Prince which Warhol used as a reference to generate an illustration. Vanity Fair then chose to buy a license Warhol's print instead of Goldsmith's photograph. So, despite the artwork being visual transformative (silkscreen vs photograph) the actual use was not transformed.
- thedevilslawyer 1y agoWhat's the steelman case that is transformative? Because prima-facie, it seems to only output original output - "intelligent" output.
- derbOac 1y agoBut those training the LLMs are still using the works, and not just to discuss them, which I think is the point of fair use doctrine. I guess I fail to see how it's any different from me using it in some other way? If I wanted to write a play very loosely inspired by Blood Meridian, it might be transformative, but that doesn't justify me pirating the book. I tend to think copyright should be extremely limited compared to what it is now, but to me the logic of this ruling is illogical other than "it's ok for a corporation to use lots of works without permission but not for an individual to use a single work without permission." Maybe if they suddenly loosened copyright enforcement for everyone I might feel differently. "Kill one man, and you are a murderer. Kill millions of men, and you are a conqueror." (An admittedly hyperbolic comparison, but similar idea.)
- rcxdude 1y ago>If I wanted to write a play very loosely inspired by Blood Meridian, it might be transformative, but that doesn't justify me pirating the book. I think that's the conclusion of the judge. If Anthropic were to buy the books and train on them, without extra permission from the authors, it would be fair use, much like if you were to be inspired by it (though in that case, it may not even count as a derivative work at all, if the relationship is sufficiently loose). But that doesn't mean they are free to pirate it either, so they are likely to be liable for that (exactly how that interpretation works with copyright law I'm not entirely sure: I know in some places that downloading stuff is less of a problem than distributing it to others because the latter is the main thing that copyright is concerned with. And AFAIK most companies doing large model training are maintaining that fair use also extends to them gathering the data in the first place). (Fair use isn't just for discussion. It covers a broad range of potential use cases, and they're not enumerated precisely in copyright law AFAIK, there's a complicated range of case law that forms the guidelines for it)
- altruios 1y agowhich AFAIN IANAL, copyright and exhaustive rights are completely different. Under copyright, once a book is purchased: that's it. Reselling the same, or transformed (re: highlighted) worked 'used' is 100% legal, as is consuming it at your discretion (in your mind {a billion times}, a fire, or (yes even) what amounts to a fancy calculator). (that's all to say copyright is dated and needs an overhaul) But that's taking a viewpoint of 'training a personal AI in your home', which isn't something that actually happens... The issue has never been the training data itself. Training an AI and 'looking at data and optimizing a (human understanding/AI understanding) function over it' are categorically the same, even if mechanically/biologically they are very different.
- ticulatedspline 1y agoDefinitely seems reasonable to say "you can train on this data but you have to have a legal copy" Personally I like to frame most AI problems by substituting a human (or humans) for the AI. Works pretty well most of the time. In this case if you hired a bunch of artists/writers that somehow had never seen a Disney movie and to train them to make crappy Disney clones you made them watch all the movies it certainly would be legal to do so but only if they had legit copies in the training room. Pirating the movies would be illegal. Though the downside is it does create a training moat. If you want to create the super-brain AI that's conversant on the corpus of copyrighted human literature you're going to need a training library worth millions
- johnnyanmac 1y agoThat's a part of the issue. I'm not sure if this has happened in visual arts, but there is in fact precedent against trying to hire a sound a like over the one you want to sound like. You can't be in talks with Scarlet Johannsen, reject her, and then hire a sound a like and say "talk like Scarlet". It's pretty clear at that point what you want but you didn't want to pay talent for it. I see elements of that here. Buying copyrighted works not to be exposed and be inspired, nor to utilize the aithor's talents, but to fuel a commercialization of sound-a-likes.
- lesuorac 1y ago> You can't be in talks with Scarlet Johannsen, reject her, and then hire a sound a like and say "talk like Scarlet" Keep in mind, the Authors in the lawsuit are not claiming the _output_ is copyright infringement so Alsup isn't deciding that.
- Dracophoenix 1y ago> but there is in fact precedent against trying to hire a sound a like over the one you want to sound like. You can't be in talks with Scarlet Johannsen, reject her, and then hire a sound a like and say "talk like Scarlet". It's pretty clear at that point what you want but you didn't want to pay talent for it. You're referencing Midler v Ford Motor Co in the 9th circuit. This case largely applies to California, not the whole nation. Even then, it would take one Supreme Court case to overturn it.
- doctorpangloss 1y agoIt’s similar to the Google Books ruling, which Google lost. Anthropic also lost. TechCrunch and others are very aspirational here.
- philipkglass 1y agoDo you mean Authors Guild, Inc. v. Google, Inc.? Google won that case: https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,_Inc https://en.wikipedia.org/wiki/Authors_Guild,_Inc._v._Google,.... Maybe there's another big Google Books lawsuit that Google ultimately lost, but I don't know which one you mean in that case.
- doctorpangloss 1y agosee, but if you ask a copyright attorney: Google lost. This is what I mean by aspirational. They won something, in very similar circumstances to Anthropic, "fair use," but everything else that made what they were doing a practical reality instead of purely theoretical required negotiation with Authors Guild, and indeed, they are not doing what they wanted to do, right? Anthropic has to go to trial still, they had to pirate the books to train, and they will not win on their right to commercialize the results of training, because neither did Google, so what good is the Fair Use ruling, besides allowing OpenAI v. NYTimes to proceed a little longer?
- dragonwriter 1y ago> Anthropic has to go to trial still, they had to pirate the books to train They did not have to, they had an alternate means available (and used it for many of the books), buying physical copies and destructively scanning them. > and they will not win on their right to commercialize the results of training That seems an unwarranted conclusion, at best. > so what good is the Fair Use ruling If nothing else, assuming the logic of the ruling is followed by the inevitable appeals court decision and becomes binding precedent, it provides a clear road to legally training LLMs on books without copyright issues (combination of "training is fair use" and "destructive scanning for storage and searchability is fair use"), even if the pirating of a subset of the source material in this case were to make Anthropic's existing products prohibited (which I think you are wrong to think is the likely outcome.)
- SoKamil 1y agoWhat if I overfit my LLM so it spits out copyrighted work with special prompting? Where to draw the line in training?
- ninetyninenine 1y agoI mean the human brain can memorize things as well and it’s not illegal. It’s only illegal if said memorized thing is distributed.
- mrguyorama 1y agoBecause humans have rights AI models do not.
- NoOn3 1y agoExactly. If someone wants to compare AI models with humans, maybe then they give AI Models the right to vote and other rights.
- ninetyninenine 1y agoThey use to say the same thing about black people.
- martin-t 1y agoHumans don't scale. LLMs do. Even if LLMs were actual human-level AI (they are not - by far), a small bunch of rich people could use them to make enormous amounts of money without putting in the enormous amounts of work humans would have to. All the while "training" (= precomputing transformations which among other things make plagiarism detection difficult) on work which took enormous amounts of human labor without compensating those workers.
- tartoran 1y agoHumans can only memorize such few texts in comparison so they'd not be scallable in the same sense LLMs are.
- 1y ago
- almatabata 1y agoIf a publisher adds a "no AI training" clause to their contracts, does this ruling render it invalid?
- bananapub 1y agowhat contract? with who? Meta at least just downloaded ENGLISH_LANGUAGUE_BOOKS_ALL_MEGATORRENT.torrent and trained on that.
- almatabata 1y agoI know, but the article mentions that a separate ruling will be made about that pirating. quote: “We will have a trial on the pirated copies used to create Anthropic’s central library and the resulting damages,” Judge Alsup wrote in the decision. “That Anthropic later bought a copy of a book it earlier stole off the internet will not absolve it of liability for theft but it may affect the extent of statutory damages.” This tells me Anthropic acquired these books legally afterwards. I was asking if during that purchase, the seller could add a no training close to the sales contract.
- shagie 1y agoWhat contracts? And would it run afoul of first sale doctrine? https://en.wikipedia.org/wiki/First-sale_doctrine https://en.wikipedia.org/wiki/First-sale_doctrine > The doctrine was first recognized by the Supreme Court of the United States in 1908 (see Bobbs-Merrill Co. v. Straus) and subsequently codified in the Copyright Act of 1909. In the Bobbs-Merrill case, the publisher, Bobbs-Merrill, had inserted a notice in its books that any retail sale at a price under $1.00 would constitute an infringement of its copyright. The defendants, who owned Macy's department store, disregarded the notice and sold the books at a lower price without Bobbs-Merrill's consent. The Supreme Court held that the exclusive statutory right to "vend" applied only to the first sale of the copyrighted work. > Today, this rule of law is codified in 17 U.S.C. § 109(a), which provides: > Notwithstanding the provisions of section 106 (3), the owner of a particular copy or phonorecord lawfully made under this title, or any person authorized by such owner, is entitled, without the authority of the copyright owner, to sell or otherwise dispose of the possession of that copy or phonorecord. --- If I buy a copy of a book, you can't limit what I can do with the book beyond what copyright restricts me.
- veggieroll 1y agoBRB, I'm going to download all the TV shows and movies to train my vision model. Just to be sure it's working properly, I have to watch some for debugging purposes.
- ncruces 1y agoYou need to buy one copy of each for the fair use to apply.
- toomuchtodo 1y agoLet everyone donate their DVDs and other physical media. You don’t need to buy it, you just need to possess the media.
- veggieroll 1y agoIndeed, I forsee a "training dataset consortium" arising out of this, whereby a bunch of companies team up to buy one copy of everything and then share it for training amongst themselves (ex. by reselling the entire library to each other for $1).
- toomuchtodo 1y agoLike an Archive? Connected to the Internet?
- veggieroll 1y agoGenius!
- ninetyninenine 1y agoAgreed. If I memorize a book and I am deployed into the world to talk about what I memorized that is not a violation of copyright. Which is reasonable logically because essentially this is what an LLM is doing.
- layer8 1y agoIt might be different if you are a commercial product which couldn’t have been created without incorporating the contents of all those books. Humans, animals, hardware and software are treated differently by law because they have different constraints and capabilities.
- ninetyninenine 1y agoBut a commercial product is reaching parity with human capability. Let's be real, Humans have special treatment (more special than animals as we can eat and slaughter animals but not other humans) because WE created the law to serve humans. So in terms of being fair across the board LLMs are no different. But there's no harm in giving ourselves special treatment.
- layer8 1y agoGenerative AIs are very different from humans because they can be copied losslessly and scaled tremendously, and also have no individual liability, nor awareness of how similar their output is to something in their training material. They are very different in constraints and capabilities from humans in all sorts of ways. For one, a human will likely never reproduce a book they read without being aware that that’s what they are doing.
- jplusequalt 1y ago>So in terms of being fair across the board LLMs are no different Why should "fair" factor into it? The LLMs are not humans, thus they have no rights, and treating them fairly shouldn't come into it. Stop anthropomorphizing linear algebra ffs.
- 1y ago
- gbacon 1y agoThe HN crowd dislikes brick-and-mortar landlords but often sides with charging rent for certain bits. Which side will prevail? Interesting excerpt: > “We will have a trial on the pirated copies used to create Anthropic’s central library and the resulting damages,” Judge Alsup wrote in the decision. “That Anthropic later bought a copy of a book it earlier stole off the internet will not absolve it of liability for theft but it may affect the extent of statutory damages.” Language of “pirated” and “theft” are from the article. If they did realize a mistake and purchased copies after the fact, why should that be insufficient?
- PunchTornado 1y agowhy would it erase the mistake? you pirated first.
- gbacon 1y agoWho is the victim, and how was that person not made whole?
- kccqzy 1y agoThe copyright holder. That person was not made whole because of the time value of money. I stole $1000 from you in January and returned it to you in June: why should you happily give me a zero interest loan?
- gbacon 1y agoNo royalty contract is getting an author a thousand bucks per sale. If you have to wildly exaggerate to make your point, then the point isn’t compelling. Books have a resale market. Every “lost sale” isn’t necessarily of a new purchase from a bookstore or Amazon. Copyright has a place. Rent-seeking authors attacking LLM owners is not a sympathetic case. Said authors are demanding to have their ideas relegated to unknown backwaters. It makes the authors worse off. It makes the community poorer. Cui bono?
- bgwalter 1y agoI have the feeling that with Alsup always the larger and more recent company wins. Google won vs. Oracle, now this. So what is he going to do about the initial copyright infringement? Will the perpetrators get the Aaron Schwartz treatment?
- bananapub 1y ago[flagged]
- AnimalMuppet 1y agoIf you're going to accuse a federal judge of corruption, you'd better have something more than a bare accusation. What is your evidence that there is corruption here, rather than just a decision that you don't like?
- NobodyNada 1y agoOne aspect of this ruling [1] that I find concerning: on pages 7 and 11-12, it concedes that the LLM does substantially "memorize" copyrighted works, but rules that this doesn't violate the author's copyright because Anthropic has server-side filtering to avoid reproducing memorized text. (Alsup compares this to Google Books, which has server-side searchable full-text copies of copyrighted books, but only allows users to access snippets in a non-infringing manner.) Does this imply that distributing open-weights models such as Llama is copyright infringement, since users can trivially run the model without output filtering to extract the memorized text? [1]: https://storage.courtlistener.com/recap/gov.uscourts.cand.434709/gov.uscourts.cand.434709.231.0_2.pdf https://storage.courtlistener.com/recap/gov.uscourts.cand.43...
- deadbabe 1y agoYou can use the copyrighted text for personal purposes.
- layer8 1y agoBut you can’t distribute it, which in the scenario mentioned in the parent’s final paragraph arguably happens.
- AnthonyMouse 1y agoYou can't distribute the copyrighted works, but that isn't inherently the same thing as the model. It's sort of like distributing a compendium of book reviews. Many of the reviews have quotes from the book. If there are thousands of reviews, you could potentially reconstruct the whole book, but that's not the point of the thing and so it makes sense for the infringing thing to be "using it to reconstruct the whole book" rather than "distributing the compendium". And then Anthropic fended off the argument that their service was intended for doing the former because they were explicitly taking measures to prevent that.
- layer8 1y agoThe premise was that the model is able to reproduce the memorized text, and that what saved Anthropic was them having server-side filtering to avoid reproducing that text. So the presumption is that without those filters, the model would be able to reproduce text substantial enough to constitute a copyright violation (otherwise they wouldn’t need the filter argument). Distributing a “machine” producing such output would constitute copyright infringement. Maybe this is a misrepresentation of the actual Anthropic case, I have no idea, but it’s the scenario I was addressing.
- josefritzishere 1y agoThe US legal systel is bending over backwards to help AI development. The arguments border on nonsense.
- shadowgovt 1y agoCan you offer some examples from this ruling? It seems pretty reasonable on a first read.
- rasz 1y agoJudge decided having an output filter on your AI makes it ok for it to contain full copy of copyrighted work. Its like saying it should be legal for me to have this Judges nudes obtained 100% illegally as long as I pixelate all the naughty bits.
- shadowgovt 1y agoFull ruling is here (https://storage.courtlistener.com/recap/gov.uscourts.cand.434709/gov.uscourts.cand.434709.231.0_2.pdf https://storage.courtlistener.com/recap/gov.uscourts.cand.43...) The analogy the judge gives is to how Google Books walked the tightrope on copyright: they maintain an archive of all the books for indexing and search purposes, and can display excerpts to help you confirm that's what you're looking for. The excerpts are constrained so you can't read the whole book by scanning the excerpts. If post-filtering the LLM signal is illegal, shouldn't Google Books archive also be illegal? If not, why not? And if you believe it should be, understand that the way precedent works, the judge won't be ruling that way without pulling some fire on themselves, because it is not the business of another case to contradict the conclusions of a previous court in a previous case. Copyright law is arbitrary and highly path-dependent because the underlying goal is forever in tension with itself, that goal being providing societal benefit by creating artificial scarcity on something that is, by its nature, not scarce at all. (Worth noting: Anthropic didn't get off scot-free. The ruling was that the created artifact, the LLM, was a fair-use product, but the way it was created was through massive piracy and Anthropic is liable for that copying).
- deepsun 1y agoOk, so I can create a website, say, the-ai-pirate-bay.com, where I stream AI-reproduced movies. They are not verbatim, so I don't infringe any copyrights.
- layer8 1y agoThey will infringe copyright as soon as they are sufficiently similar to the original. You can’t shoot a non-verbatim but clearly recognizable beat-by-beat remake of Star Wars, call it Galaxy Conflict, and get away with monetizing it.
- shadowgovt 1y agoCorrect. You have to call it "Starcrash" (https://www.imdb.com/title/tt0079946/?ref_=ls_t_8 https://www.imdb.com/title/tt0079946/?ref_=ls_t_8). Then it's legal.
- layer8 1y agoInteresting artifact, but the very first/top IMDB user review convincingly contradicts that this is a Star Wars remake. ;)
- paxys 1y agoWill be interesting to see how this affects Anthropic's ongoing lawsuit with Reddit, or all the different media publishing ones flying around. Is it okay to train on books but not online posts and articles? Why the distinction between the two?
- cyanmagenta 1y agoThe distinction will be whether those online posts were obtained legally, analogous to whether the books in this case were pirated. It’s not as simple as it sounds, since I’m sure scraping is against Reddit’s terms and conditions, but if those posts are made publicly available without the scraper actually agreeing to anything, is that a valid breach of contract? Will be interesting to see how that plays out.
- bradley13 1y agoGood. Reading books is legal. If I own a book and feed it to a program I wrote (and I have done exactly that), it is also legal. There is zero reason this should be any different with an AI.
- thinkingtoilet 1y agoIf you charge me to use your program and it spits out unedited, copyrighted material then it should be illegal. I don't know the details of this case, but that's what's going on in the New York Times case. It's not always so cut and dry.
- dmix 1y agoWhich is amusing because NYTimes has fought in court a few times in favour of technology progress over copyright. Including recently when they got sued over collected a bunch of freelance writing into a database without consent. https://harvardlawreview.org/blog/2024/04/nyt-v-openai-the-timess-about-face/ https://harvardlawreview.org/blog/2024/04/nyt-v-openai-the-t... I doubt the exact replica stuff will stand, as technically it was only achievable via advanced prompt engineering (hacking), not simply asking for a replica. So their 2 other arguments boils down to scraping a news database = infringement and LLM output = derivative works.
- deleted 1y ago[deleted]
- szc 1y agoI've co-authored a book that a lot of the models seem to know about. The models consistently get the names of the authors incorrect and quote the material with errors. If the canonical representation of our work is now embedded within AI models, don't we deserve to have it quoted and represented correctly and fairly? If you asked a human who had read the book, I think there is a fair chance they would likely give you the reference to the source material. I do concede that the book does contain a distillation of material that is also available from other sources, but it also contained a lot of personal experience. That aspect does seem to be lost in this new representation. I am not saying that letting AI models read the material is wrong, but the hubris in the way models answer questions is annoying.
- UltraSane 1y agoIf the US makes it illegal to train LLMs on copyrighted data that isn't going to stop China from doing it and give them an ENORMOUS advantage.
- rsstack 1y agohttps://news.ycombinator.com/item?id=44369227 https://news.ycombinator.com/item?id=44369227 If the US makes it illegal to train LLMs on copyrighted data, the US will find a solution and not just give up and wait half a decade to see what China does in the meantime.
- UltraSane 1y agoWhat solution is there?
- rsstack 1y agoZillow have the MLSs network that provide them lists, a similar solution could apply if courts agree that library copies apply for this - Anthropic could sign agreements with large libraries and "check out"/"freeze" copies for a minimally-agreed-upon duration and query across all to see which has a copy of each book they need. Spotify and Apple Music sign deals en masse with labels, the same could apply here with book publishers, labels for lyrics, museums for art, etc. Or whatever other creative solution that people who will need to find, will find. Right now they took the laziest path, because it worked. They will find the next-laziest path that works. And the easiest option: Legislation change. If it's completely decided that the current law blocks LLMs from working in the US, the industry will lobby to amend the copyright law (which is not immutable) to add a carveout for it. You're assuming that people will just give up. People never gave up, why would they now?
- kmeisthax 1y ago> “We will have a trial on the pirated copies used to create Anthropic’s central library and the resulting damages,” Judge Alsup wrote in the decision. “That Anthropic later bought a copy of a book it earlier stole off the internet will not absolve it of liability for theft but it may affect the extent of statutory damages.” I'm not sure why this alone is considered a separate issue from training the AI with books. Buying a copy of a copyrighted work doesn't inherently convey 'fair use rights' to the purchaser. If I buy a work, read it, sell it, and then publish a review or parody of it, I don't infringe copyright. Why does mere possession of an unauthorized copy create a separate triable matter before the court? Keep in mind, you can legally engineer EULAs in such a way that merely purchasing the work surrenders all of your fair use rights. So this could wind up being effectively: "AI training is fair use for works purchased before June 24th, 2025, everything after is forbidden, here's your brand new moat OpenAI"
- comex 1y agoThe ruling suggests that "pirating a book that could have been bought at a bookstore" for the sake of "writing a book review" "is inherently, irredeemably infringing". Which suggests that, at least in the judge's opinion, 'fair use rights' do exist in a sense, but it's about when you read the book, not when you publish. But that's not settled precedent. Meta is currently arguing the opposite in Kadrey v. Meta: they're claiming that they can get away with torrenting training material as long as they only leech (download) and don't seed (upload), because, although the act of downloading (copying) is generally infringement under a Ninth Circuit precedent, they were making a fair use. As for EULAs, that might be true for e-books, but publishers can't really do anything about Anthropic's new strategy of scanning physical books, because physical books generally don't come with shrinkwrap license agreements. Perhaps publishers could start adding them, but I think that would sit poorly with the public and the courts. (That's assuming the ruling isn't overturned on appeal, which it easily might be.)
- kmeisthax 1y ago[dead]
- 1y ago
- Fluorescence 1y agoI'm surprised we never discuss a previous case of how governments handled a valuable new technology that challenged creative's ability to monetise their work: Cassette Tapes and Private Copying Levy. https://en.wikipedia.org/wiki/Private_copying_levy https://en.wikipedia.org/wiki/Private_copying_levy Governments didn't ban tapes but taxed them and fed the proceeds back into the royalty system. An equivalent for books might be an LLM tax funding a negative tax rate for sold books e.g. earn $5 and the gov tops it up. Can't imagine how to ensure it was fair though. Alternatively, might be an interesting math problem to calculate royalties for the training data used in each user request!
- bonoboTP 1y agoSurely this would require the observation that the public is actually using LLMs as a substitute for purchasing the book, ie they sit down and type "Generate me the first/second/third chapter of The Da Vinci Code" and then read if from there. Because it was easy to observe in the cassette tape era that people copied the store bought music and films and shared it among each other. I doubt that this is or will be a serious use case of LLMs.
- munificent 1y agoFor narrative fiction, probably not, no. But for non-fiction and technical books, they sure as hell will. For example, if I ask Gemini "What does Bob Nystrom say about the Singleton pattern in Game Programming Patterns?", it says: --- In "Game Programming Patterns," Bob Nystrom dedicates an entire chapter to the Singleton pattern, which is unusual because, as he states, "Every other chapter in this book shows you how to use a design pattern. This chapter shows you how not to use one." Here's a summary of his key points regarding the Singleton pattern: * More Harm Than Good: Nystrom argues that despite its noble intentions, the Singleton pattern often does more harm than good, especially in game development. He notes that the Gang of Four (GoF) stressed its sparing use, a message that was often lost in translation to the game industry. * Global State in Disguise: He highlights that the Singleton is essentially a global variable in disguise. While it ensures a single instance, it often leads to problems associated with global state, such as tightly coupled code, difficulty in testing, and reduced flexibility. * Encourages Tight Coupling: By providing a global access point, Singletons encourage various parts of the codebase to directly depend on them. This makes the code harder to change and refactor because modifications to the Singleton can have widespread, unpredictable effects. * Hinders Testability: Testing code that relies on Singletons can be challenging. Because Singletons are globally accessible and often difficult to mock or replace, it can be hard to isolate units of code for testing, leading to brittle tests or a lack of testing altogether. * Limits Flexibility: Nystrom points out that while a class might currently only need one instance (e.g., a FileSystem or AudioSystem), hardcoding this limitation with a Singleton makes it difficult to change in the future. If a game later needs multiple file systems (e.g., for different save slots) or different audio contexts, a Singleton makes this a non-trivial refactoring. * Alternatives Often Better: He suggests that often, what developers think they need a Singleton for can be better achieved with other patterns or simpler solutions. For instance, if the only goal is to have a single instance, a simple static class might suffice if global access is truly desired and initialization order isn't a complex concern. For managing services, he often advocates for the Service Locator pattern as a less problematic alternative that still provides global access without enforcing a single instance. * Misconceptions about "Singular Resources": While the GoF suggested Singletons for "physical devices," Nystrom argues that few physical devices are truly singular in a way that necessitates a Singleton. For example, while you might only have one keyboard, a game could theoretically support multiple input devices. Using a Singleton prematurely prevents this flexibility. * Lazy Initialization: He acknowledges that a well-implemented Singleton using lazy initialization (creating the instance only when it's first accessed) can save memory and CPU cycles and resolve some issues with static initialization order. However, he emphasizes that these benefits don't outweigh the pattern's fundamental architectural drawbacks. * In essence, Nystrom's strong stance is that while the Singleton pattern can ensure a single instance and provide global access, its downsides (coupling, testability, flexibility) often make it an anti-pattern in game development, and developers should carefully consider alternatives before reaching for it. --- Is that summary as good as actually reading the book? Probably not. Will people rely on that and skip buying my book. Almost certainly so.
- blindriver 1y agoHumans read books. AI/LLMs do not read. I think there's an inherent difference here. If the LLM is making a copy of the entire book in it's memory, is that copyright infringement? I don't know the answer to that, but it feels like Alsup is considering this fair use argument in the context of a human, but it's nothing like a human and needs to be treated differently.
- steveklabnik 1y agoLLMs do not "make a copy of the entire book in its memory" so that specific question is kind of moot.
- rasz 1y agoIts already established it can recite whole Hairy Potter and Carmacks Fast Inverse word for word. Just because it uses fancy compression doesnt mean its not a copy.
- deleted 1y ago[deleted]
- riskable 1y agoIt can recite something like 80% of Harry Potter with carefully crafted prompts. If you take half a sentence from Harry Potter then tell the LLM to predict the rest it will complete it. That's what they did in that study you're referring to. It's not even remotely the same thing as "can recite whole Harry Potter." If you ask an LLM to regurgitate Harry Potter it won't be able to do so because that's not how they work. They're prediction engines and it just so happens that Harry Potter quotes/excerpts are so pervasive on the Internet that the LLMs ingress ranks that style of wording higher than other styles. Ask it to regurgitate some other, less-popular work. Do it for hundreds or thousands of them. You'll quickly find that those two examples you gave are the exceptions and that LLMs can't pull it off. They won't even get close.
- rasz 1y ago
- nektro 1y agodevastating news
- sillysaurusx 1y agoThe reason I made books3 was to help force a decision on this issue. I’m happy to see that it’s settled, and that it’s legal for robots to read books. It’s also proof that an individual scientist can still change the world, in some small way. Believe in yourself and just focus on your work, even if the work is controversial. (I’m late to the thread, so ~nobody will see this. But it’s the culmination of about five years of work for me, so I wanted to post a small celebratory comment anyway. Thank you to everyone who was supportive, and who kept an open mind. Lots of people chose to throw verbal harassment my way, even offline, but the HN community has always been nice.)
- mgr86 1y agoFWIW, I see your comment. Also late to the thread though. This ruling is being watched at my office. I want to be a bit anonymous, but we've been doing a much more analogue version of some of these things for 75 years. With academics being our primary market. We've only had two legal issues in that time. Both settled out of court. But we walk a fine line.
- sillysaurusx 1y agoThank you for your work!
- jplusequalt 1y agoI think you have indirectly done a disservice to the artistic community.
- lofaszvanitt 1y agoOf course, it's the United Steal of America.