11 ms·
I dislike these AI companies but let's be clear here: the copyright holders are the ones locking these books up. If they don't want to print more copies, then t
by ezfe 1mo ago
I dislike these AI companies but let's be clear here: the copyright holders are the ones locking these books up. If they don't want to print more copies, then they could release the copyright on them.
Instead, they enforce the copyright and force AI companies to shred books they want to ingest.
edit: Also, an AI company would only ever care to purchase, scan, destroy a book once. Presumably many books have more than one copy.
- breezybottom 1mo agoThey don't "force" anything. Trillion dollar AI companies and their owners have as much agency as book publishers.
- tptacek 1mo agoTo do what?
- deleted 1mo ago[deleted]
- runarberg 1mo agoTo not destroy rare books.
- tptacek 1mo agoWhen you read "rare books", what are you thinking these are? The 404 article that spun this story up goes into more detail. These are like vanity press books. They're rare because nobody cares about them. The book industry already destroys these books.
- runarberg 1mo agoI am thinking about a long essay Icelandic author Þórbergur Þórðarson wrote to his pals abroad, and were left abroad. I am thinking about a photo book by an Indonesian naturalist who is famous on Bali, and took amazing photos of wildlife on Sulawesi in 1926, and colored in, and somehow ended up in New Jersey in the 1980s. I am thinking about a collection of essays written by a teenage J.D. Salinger who he left unsigned at a café thinking nobody would want to read them but just leaving it up to chance. Or maybe a Jackson Pollock sketchbook he lost at a party which ended in the host’s bookshelf, and finally at an estate sale. Plenty of such unknown unknown exist, and the AI machine will inevitably destroy a bunch of them at this scale.
- tptacek 1mo agoYou're trying to imagine rare valuable books and then fantasizing about AI companies destroying them, but what's actually happening here is that AI companies are acquiring, digitizing, and then pulping the instruction manuals to 1983-vintage copy machines. This is all such a special-pleading argument. You know what other institution snatches up books and destroys them at huge scale? Public library systems. People clean out their attics and basements and drop off huge boxes full of books at libraries; libraries take the things they know will circulate, and destroy the rest. Take a guess as to how Þórbergur Þórðarson fares at the Newark Public Library. Wait, bad example, they stopped accepting book donations because nobody wants your old books. They tell you to give the books to thrift stores instead. Guess what the thrift stores do with them? You know how many times I've read stories about the grave damage libraries are doing to human culture? Zero, zero times.
- frm88 1mo agoYou're trying to imagine rare valuable books and then fantasizing about AI companies destroying them, but what's actually happening here is that AI companies are acquiring, digitizing, and then pulping the instruction manuals to 1983-vintage copy machines. Source? You state that in a tone that implies you have verifiable knowlege of this. All the information I found says that the exact number, titles and authors are under NDA.
- fenomas 1mo agoNot under US copyright law. The Bartz case ruled that if you scan a book and destroy the original it's considered format-shifting and you're likely fine. But not so for keeping the original and using the scan in its place - Internet Archive tried that (in an incredibly limited way), and publishers sued and won. So companies scanning books already know they'll be sued, successfully, if they don't destroy the originals. So they destroy the originals.
- parineum 1mo agoThey are "forced" to do this because that's what they have to do to abide by copyright law. They can't create a digital duplicate without destroying the original.
- breezybottom 1mo agoSince when do AI companies care about copyright law? They're destroying them so their competitors can't use them.
- alightsoul 1mo agoSince they had to pay 1.5 billion for it
- freejazz 1mo agoYou'd think they could've negotiated something better than $3k per work if they had actually gone to the table first.
- eru 1mo ago> They're destroying them so their competitors can't use them. That's pretty silly. My competitor can't ride my bike either, and I didn't have to destroy the bike for that to be true.
- gpt5 1mo agoA little meta - I want to point out demagogic/populist comments like these that try to clear all nuance and brush a topic in black and white tend to come from a really small portion of the users here, but the same user (whose account is only 4 months old), dominates posts like this by posting many many bait-like comments in the same post that deviate the discussion away from meaning and insight. This is just one example, but it has become unfortunately common across all social media platforms.
- freejazz 1mo agoIf that's true, then why did they pirate so many books?
- zmmmmm 1mo agoIt is also the case that the copyright holders are often putting restrictions around use of electronic forms that are driving the desire to use physical copies. I doubt AI companies would use a single physical book if they could avoid it - absent the legal cloud over electronic rights. I have no evidence but I can't help suspecting in part the publicity around this is driven in part by rights holders that want to force AI companies back to e-books where they can force them into licensing deals.
- alightsoul 1mo agoAi companies don't use ebooks, because they are more expensive than second hand books
- breezybottom 1mo agoThey absolutely do. Meta torrented 81 terabytes of ebooks. They just have no incentive to pay when the law looks the other way.
- alightsoul 1mo agoI meant paid ebooks. That's probably what the commenter refers to, because that's what publishers want. Obviously ai companies don't want to pay so they try to use pirated ebooks
- hn_throwaway_99 1mo agoThe entire ironic thing here is that a huge part of those 81 terabytes of ebooks that Meta torrented were directly pirated books from Anna's Archive.
- bawolff 1mo agoi imagine its because the doctrine of first sale does not apply to ebooks.
- warkdarrior 1mo agoOn Amazon right now, retail prices for e-book copies are higher than for the corresponding paperbacks.
- jacobo37 1mo agothis is plainly stupid ... many of these books are likely to have no current publisher nor any way to "reprint" the book. "ai" companies are simply burning our cultural context ...
- rpdillon 1mo agoWait: the entire premise of copyright is to prevent someone from publishing a book, and a competitor buys a copy, clones it, and sells copies way cheaper because they don't have to pay the author. Now, in 2026, we're acting like cloning a published book is not technically feasible? That doesn't track. With publishing on-demand, it's easy to imagine a business with digital copies of all these works that they make available for print-on-demand. The uncomfortable reality is that most of these books are nothing anyone cares about. Even the book sellers in the 404 story call them dead inventory. Can we get some actual book titles into the discussion so we can focus on facts rather than speculation?
- alightsoul 1mo agoThis is not a technical problem at all. This is a copyright problem. Anthropic thought it was just a technical problem until they had to pay 1.5 billion after they lost a copyright court case Op means a lot of those books were made before computers were used for that purpose and the publishers and probably authors no longer exist, so there is no digital copy to just reprint, unless someone scans it themselves and publishes it, risking copyright violation when done at large scale due to possible exceptions to this rule
- Sha1rholder 1mo ago> Anthropic thought it was just a technical problem Did they? Then why did they "don't want anyone to know about this"? Or do you think their lawyers are dumb?
- dukeyukey 1mo agoI imagine they know the public would react like the public is reacting. Like obviously this is not worse than what libraries and thrift shops do daily, but it is bad optics.
- HedonicEscal8r 1mo agoIf only this complaint was being posted by an organization ideologically opposed to copyright itself!
- RajT88 1mo agoThe articles I've read on this are not clear, but I strongly suspect "rare" is not the definition you and I probably use for the level of rarity of books actually being destroyed. These are not going to be the kinds of books "The Ninth Gate" resolved around - truly one of a kind. It's not good they are destroying books, but they are books which do have other copies. Just perhaps not many.
- runarberg 1mo agoAt this scale, there are no guarantees of anything. There very likely will be unique copies in there. If these were expert archivists a lot of damage could be prevented, but given the malice and indifference of AI companies, there very likely will not be an expert archivist involved, and unique copies will be destroyed unceremoniously.
- eru 1mo agoMy personal wastebook at home is so rare, it's unique. That doesn't mean it needs preservation.
- runarberg 1mo agoLike I said, at this scale, there are no guarantees for anything. Very likely will there be a unique copy of an invaluable book or letter an the person feeding it to the scanner will not know and the book get destroyed. Like did Icelandic author Þórbergur Þórðarson ever write an a book about Esperanto, and send the only copy of it to Halldór Laxness when he was in Los Angeles? I don‘t know, but it is certainly something he is likely to have done. If such a book exists it would be invaluable to both Icelandic culture and to Esprentists. It likely would have stayed in Los Angeles where nobody would know the significance of it until it ended up in an estate sale, a used book store, and then finally destroyed by an AI company never to be discovered. My hypothetical is just one of trillions of possibilities. At this scale very likely several of these possibilities will unessiseraly remain unknown unknowns forever.
- 1mo ago
- deleted 1mo ago[deleted]
- raincole 1mo ago> Instead, they enforce the copyright and force AI companies to shred books they want to ingest. What? Even if there are no copyright holders, the AI companies will still do scan'n'destroy because it's just cheap. Are you expecting the authors/publishers to send digital copies to AI companies directly? Or expecting AI companies to preserve the physical copies indefinitely? Both are not gonna happen, copyrighted or not.
- remus 1mo ago> ...AI companies will still do scan'n'destroy because it's just cheap. There is also a legal element. If they kept the physical copy around after scanning the argument is that they're making copies of the book which puts them on tricky legal ground. By destroying the physical copy they can argue that there is only one version of the book that now exists solely in digital form, so this usage is better protected under fair use.
- cm2012 1mo agoYes. I dont understand at all what AA is worried about. One copy of a book is no big deal? good will and used book stores throw out a lot more than that.
- wesleywt 1mo agoThey are scanning "rare" books. I presume there are not a lot of copies left to throw out.
- joshstrange 1mo agoRare by whose definition? I’m not aiming this at you directly by: ISBNs or STFU Show me which “rare” books they are destroying and _maybe_ I’ll care but so far the pearl-clutching over this leads me to believe it’s people worked up about the idea of destroying (except it’s not destroying, it’s transforming, a fact often ignored) books, books that it’s not clear at all there is any strong demand for. People want to invoke things like F451 but it doesn’t compare in the slightest. It’s like when people get mad about libraries throwing away or otherwise liquidating books that no one is reading in order to bring in books people want to read. People get all up in arms about that as if a book itself, in isolation, is inherently valuable or worth protecting. It’s not. If no one wants to read it then what value does it have? The impetus is on the people that think the book has value, it’s on them to carry the torch, to preserve what they think is worthy. It would be like a company going to a yard sale and buying unsold/unwanted items to 3D scan them and destroy them in the process. This isn’t breaking into the Louvre and destroying one-of-a-kind artwork.
- squidbeak 1mo ago> Show me which “rare” books they are destroying and _maybe_ I’ll care but so far the pearl-clutching over this BBC good enough for you? https://www.bbc.com/news/articles/cp3rprx2wl4o https://www.bbc.com/news/articles/cp3rprx2wl4o "A recent academic text published in only 100 copies, 75 of which are already in libraries, may be very rare on the market - but it is perhaps not such a great loss if one copy is destroyed," says Derek Walker, owner of Edinburgh bookshop McNaughtan's. "But we have, and have sold, books which are for example the only known surviving example of an edition from the 18th century. "It would be a much more significant problem if one like that were to be bought for destruction, having survived this long."
- ls-a 1mo ago[dead]
- sophacles 1mo agoIf you're buying second hadn books by the lot, you'll get a lot of duplicates and its eaiser to scan wholesale and dedupe in the computers than it is to try to run a sorting operataion on "things".
- jscd 1mo agoSorry, is your stance seriously that authors and publishers should digitize and freely distribute their work, at their own expense? Also, who’s forcing AI companies to “ingest” books in such a destructive way? Also also, if there’s one thing I’ve learned from AI scrapers, it’s that they’d never scan the exact same thing multiple times at the expense of public access to the resource.
- scarmig 1mo agoRelinquishing copyright does not imply any of the labor you're suggesting. It's the opposite: you're just committing not to perform the labor of pursuing legal action against someone who does digitize and freely distribute the work. Anna's Archive, for one, would be more than happy to host at no cost to the author.
- jeroenhd 1mo agoAI companies are buying the physical books, they can turn them into confetti if that's what they want to do. If the physical books are running out, the authors can print and sell more. Or they can sell digital copies so the information is not lost. The law is currently forcing these companies to destroy the books after scanning them.
- jscd 1mo ago> The law is currently forcing these companies to destroy the books after scanning them. No, the law is stopping them from digitally sharing their scans. They are perfectly capable of reselling or donating or storing the books they buy. (Wasn’t Amazon originally a book seller?)
- ezfe 1mo ago> Sorry, is your stance seriously that authors and publishers should digitize and freely distribute their work, at their own expense? digitize: no, there's no burden to do this freely distribute: no, there's no burden to do this use it or lose it on the copyright: yes; if a work is copyright but a publisher doesn't want to make new copies because they won't make money on it then it should be free to copy
- wotamess 1mo ago"want to ingest" Not "need to ingest" Copyright holders are capitalizing on laws on the books just like Jeff Bezos companies buying their own copies to shred So in the end it's really a Congress problem as usual
- customguy 1mo ago> force AI companies to shred books they want to ingest. Nothing forces them to shred books, they do it because it's slightly cheaper that way.
- postepowanieadm 1mo agoBy destroying them they don't copy only convert them into another format.
- hparadiz 1mo agoThere was a court case where they said that if they copied the books it's not fair use because they didn't own it but if they bought physical copies and then destroyed them then somehow it was fair use because it fell into the niche of personal backups. I forget the details but basically they buy one time prints and destroy them immediately.
- customguy 1mo agoI had no idea about that, or how fucking bad this actually is: https://en.wikipedia.org/wiki/Project_Panama https://en.wikipedia.org/wiki/Project_Panama So I stand corrected: at least some don't do it because it's cheaper (than to buy a license, or simply forego some things), but because they're fucking evil, or so stupid that it effectively is the same as being extremely evil.
- skeledrew 1mo agoThey do it because, for each work, they bought one copy, which they scan and no longer need the physical version of, and would be in copyright violation if they keep more copies than they bought.
- annapanna 1mo ago>and would be in copyright violation if they keep more copies than they bought. They can contact the copyright holder and ask/buy a license to make multiple copies.
- ErigmolCt 1mo agoCopyright holders are certainly responsible for keeping unavailable works inaccessible, but shredding is mostly an industrial scanning decision, not a copyright requirement
- ezfe 1mo agoIt’s the outcome of a US court case that requires destruction of the original work
- wesleywt 1mo agoI was wondering what the pro book shredding take was going to be. Why destroy the book after scanning? You can create a beautiful library of rare books with all the AI debt bubble.
- watwut 1mo agoCopyright allows you to sell book you have and does not force you to shread it. They are not forced to shread them by copyright.
- signa11 1mo agomr. vernor-vinge's "Rainbows End" is oddly prescient ! Highly recommended nevertheless.
- bondarchuk 1mo agoWe the people in the society who have the power to make laws through democratic means are the ones locking these books up.
- ctm92 1mo agoThey also only scan books that are easily and cheaply available, which means they are either not rare or have no significance. Books that are rare of have historic significance will surely be in museums or libraries and not going away for pennies.
- ajsnigrutin 1mo agoIt's also the regulation, where most systems still look at "one pirate copy" = "one sale of lost profits", especially when pirates end up in court. If the book (or game or whatever) is not sold anymore in any way where you could give the copyright holder money in an easy accessible way (eg. buy it on amazon, or a local bookstore), they shouldn't be able to claim losses from piracy, since they clearly don't want your money. On the other hand, there are grey zones here, the lord of the rings books (still copyrighted and easily obtained pretty much everywhere) have been translated into my language many decades ago, and many of us read and liked those translations, but when the movies came out, a new translator did a new translation, where they changed a lot of things, including the last names of bilbo and frodo (Bogataj->Bisagin) and the Shire (Grofija->Šajerska), and the old version is sadly available only in paper form on second hand markets. On one hand, copying that if you only want this specific version would not cause a lost sale, on the other, you can get new translations (or english originals) pretty much everywhere.
- asdefghyk 1mo agoRE "...Instead, they enforce the copyright and force AI companies to shred books they want to ingest...." Why are AI companies forced to shred books?
- jeroenhd 1mo agoThey want to have a digital copy and the judge ruled they can only keep one copy.
- drtgh 1mo agoYou do not need to shred books to scan them. This only happens if you don't care about preserving the integrity of the books and you want to scan more cheaply.
- jeroenhd 1mo agoBut their legal framework for being permitted to scan the books en masse (they are "transforming" the book from physical to digital) requires destruction of the original. Otherwise it wouldn't be transforming, it would be duplicating.
- deleted 1mo ago[deleted]
- asdefghyk 1mo agoOK that action allows them to do the transformation. But how about the importing into the AI tool? Does this "transformation" somehow override the authors complaints of AI taking their book and not compensating the authors for its mass usage in the AI tool.
- drtgh 1mo agoThen, should we hope they do not do backups, as they would be duplicating. And that they delete the files (zeroing from the disk) when such books are loaded into memory as technically it would be duplicating also. The same if they use different machines simultaneously.
- ranit 1mo ago> Also, an AI company would only ever care to purchase, scan, destroy a book once. Presumably many books have more than one copy. The OP article sounds quite opposite though - that AI companies are doing exactly this - destroying books so only they have the scanned content.
- tgsovlerkhgsel 1mo agoNot the copyright holders, "we the people": Copyright is an artificial legal construct that was repeatedly ratcheted up over and over again. Unfortunately 50 years after the death of the author (or 50 years after publication for corporate owned works) has been locked in as a minimum term through international treaties, so it'll be somewhat hard to lower it beyond that, but many countries (including the US) enforce much longer terms, so that would be a first lever that could be applied quickly. Maybe countries could could also establish an exception for out of print books offered to the public for free, or a general "library exemption" for public archives after a certain number of years? I'm sure one of the AI companies would be willing to host a LibGen style library as a PR measure if legally allowed (with sign up required for rate limiting and as an extra benefit for the company to get daily active users).
- lkbm 1mo agoWith many old books, a big part of the problem is that it's non-trivial to determine who owns the copyright. Sometimes the contract would say the copyright reverts to the author after a certain amount of time out of print,but you have to go dig through old contracts to figure out whether that's the case for any given book.