7 ms·
I was surprised to read that Anthropic (and probably other data / model companies) are doing this and it's extremely disappointing, as working towards the benef
by ziyadb 1mo ago
I was surprised to read that Anthropic (and probably other data / model companies) are doing this and it's extremely disappointing, as working towards the benefit of humanity is not an exclusive right / domain of theirs but rather, is a shared responsibility and mission carried by all of humanity itself as a collective responsibility. Thus, the preservation of this knowledge, its availability, and accessibility are the most important things that we must ensure continue.
From a historical precedent standpoint, this is akin to the burning of the library of Alexandria, where centuries of knowledge was destroyed and leaving a limited version of the history, the surviving one, and depriving successive generations of significant amount of latent knowledge.
Despite the copyright restrictions that are forcing companies to do this, they should maintain archives that are publicly available. As stated in my first paragraph, they are not the exclusive stewards of humanity despite them anointing themselves as such. Granted, a lot of these books might not be that useful, but still a relic of times pre-machine generated text, which makes them valuable if only for their archival value.
- smalltorch 1mo agoSurely they have the high quality scans, but there would probably be the same legal restrictions to just share the archive.
- Filligree 1mo agoObviously. Copyright infringement is settled law.
- merely-unlikely 1mo agoCopyright infringement has a long and deep bank of caselaw but as Anthropic has already discovered, it is not entirely "settled."
- JKCalhoun 1mo agoI'm only one person, but I scan old books that had an impact on me growing up, and upload them to archive.org. Thankfully there are others that do the same. (And to be sure, FWIW, these are books that have not been printed for about 50 years—I suppose the software community would call them abandonware.)
- pessimizer 1mo agoIf they're 50 years old they're young, and archive.org will likely block access. If they're not already on annas-archive (or the copy there is trash), your best bet is an anon upload to libgen.
- JKCalhoun 1mo agoThanks. There was a time of course when you could pull my books down from archive.org as PDFs. Perhaps that time will come again. I'll look into libgen.
- ceasesurthinko 1mo agoArchive.org Scanned Book Downloader Bookmarklet https://gist.github.com/cemerson/043d3b455317d762bb1378aeac3679f3 https://gist.github.com/cemerson/043d3b455317d762bb1378aeac3...
- mbeavitt 1mo agoFrom their perspective, it's training data that their competitors don't have. If they make it available, they fill in their moat.
- silverwind 1mo agoMore importantly: Once Anthropic is gone, all knowlege is lost.
- psma_egeliaa 1mo agoIt will probably be actioned off in the bankruptcy proceedings.
- jonhohle 1mo agoThat’s an interesting angle. There’s probably some property value (though maybe not enough based on volume) to the books they purchased. I doubt there’s any value to the “backups” of those books. I’d imagine they’re normally transferable.
- wildzzz 1mo agoThey cut the bindings off and feed loose leaf books through a document scanner. It would be difficult to store and probably unsellable, it's probably going in the trash. To legally retain these scans, you must own the original book. You can sell or give away the original book (sans binding) but it's legally dubious as to whether the scan can be transferred along with it. So if Anthropic has no interest in storing thousands of loose leaf books, they are likely destroying both the original and scan as soon as possible. At the end of the day, the only thing of value Anthropic has is the trained model which is definitely transferable.
- merely-unlikely 1mo agoIt's already been distilled.
- flatline 1mo ago
- alerighi 1mo agoYes but let's continue using Claude to write code because we suck at programming. Really the only way out of this is to STOP NOW using AI and use our brain instead. These company will just shut down if we stop using, and thus paying, for their services. Come on, we did without AI for all our history, we could live without it with no issue (as to me we could live without smartphones, internet, etc if we want).
- HeWhoLurksLate 1mo agowe also did without air conditioning, plumbing, democracy, and human rights for millenia, and I wouldn't want to give any of those up
- yehat 1mo agoNobody will ask you, they'll be taken from you, in case you missed what happens around.
- JKCalhoun 1mo agoIf you are suggesting that air conditioning, plumbing, democracy, and human rights are bad for society then I am missing the analogy.
- alerighi 1mo agoI wouldn't even compare AI to any of one. Rather, AI will even contrast with some of these, democracy for example, if we stop critically thinking and delegate everything to AI (= big corporations that control it) it cannot end well to me. Same thing if we continue building datacenters that generate tons of CO2 for doing stuff we can do with well, our brain (that is the most energy efficient computer on earth!) or even better not do at all (because we don't need AI slop), air conditioning will not be enough I fear. Computers and computer programs did work much more reliably when they were written by humans, a modern computer system written mostly with AI has bugs that not even in systems that were in use in the 80s (and some of them are even still in service today!) had. Do you want a planet where the human being is at the center, or the machine? Do you think machines has to be at your service, that you are the one ordering it what to do (that is programming it), or you think you are to be a servant of an AI that takes all the decision for you? Because if you think the second, you want MATRIX or SKYNET, well I don't want MATRIX or SKYNET to be fair (not that it's even possible btw, since AI has anything intelligent in it beside the name, it's just a glorious copy/paste machine).
- raptor99 1mo agoI hate to be the bearer of bad news but you really do have to assume the worst about any of these "AI" companies, especially the large ones like ChatGPT and Anthropic. They literally lie, cheat and steal at any opportunity they have and in any way that they think of. Do not trust a single thing that they say; it is a fool's folly to do so. A lot of this can already be said about a lot of companies, especially almost any large company, but it goes doubly if not triply so for this new breed of company now.
- brookst 1mo agoAre you really advocating for assuming things with no evidence, by presenting no evidence for why one should do so? That’s not especially rigorous thinking.
- JKCalhoun 1mo agoIt's likely more along the "fool me twice" category of wisdom. The opposite would instead be a kind of naive thinking.
- SkyBelow 1mo ago>Despite the copyright restrictions that are forcing companies to do this, they should maintain archives that are publicly available. Aren't the copyright laws forcing them to do this the very ones that would make such archives illegal? The books that could be in such an archive are the books that don't need to be destroyed.
- JKCalhoun 1mo agoCopyright law eating itself…
- convolvatron 1mo agono matter how you look at it, this is a systemic failure. if as a society we're going to mass scan our history then we should be building an archive for the future. not using availability of information as a moat. not doing it over and over again and throwing it away because of some odd rules to protect someones market position. not using it as an excuse to put paywalls around 80 year old field guides to field rodents in western massachusetts. not taking texts that had limited value and mining them for turns of phrase to be piled up into a useless grey goo.
- hypendev 1mo ago>From a historical precedent standpoint, this is akin to the burning of the library of Alexandria, where centuries of knowledge was destroyed and leaving a limited version of the history, the surviving one, and depriving successive generations of significant amount of latent knowledge. How tho? Ever sunday at the flea market, I see thousands of books that are rotting, hoping for someone to buy them or at least take them home, so the seller doesn't have to pack them for the trip back. Just the other day, there was a whole bin of books in front of a shop, offering them for 50 cents a piece. They will be destroyed anyways. Unless they are buying and destroying really old, rare books or important small-print books, it is not much damage. It is not like they will buy "all copies of all of the books", just one. And its just that the data in physical print most likely hasn't been used for training, so this can help you find more unmined quality data. Nobody is stealing your books, preventing you from buying more or destroying all copies of a single book. And some of these books would rot out of circulation or be destroyed anyways. Some people throw away 80-100 year old books on the regular, as they might just be unimportant to them or the world in general. And once the last copy is thrown or rots, that book will die forever. This way, it will live forever instead, scanned and trained on, conjoined with the rest of our knowledge in a magic machine.
- starkd 1mo agoGood to see someone making this point. I'm confused by the panic, because they are making it out like AI companies are destroying every copy of the book. They only need one, and they destroy it after scanning it only because they don't want to store them all. And storing or archiving all these books is not a trivial task.
- sergimansilla 1mo agoNo, they are destroying them because that’s the way that they can keep a digital copy legally.
- beering 1mo agoThis is not true. Google Books does not destroy books. No court has ruled that you must destroy books to legally keep a digital copy.
- leonidasrup 1mo agoA small change in the copyright law would fix this problem. Something like: If a company is scanning material protected by copyright, it has to send a digital copy of the scanned material to Library of Congress within 5 working days.
- Upvoter33 1mo agoI love this idea. At least make them turn it into some form of a public good.
- starkd 1mo agoSo now the Library of Congress has to manage all these submissions whenever someone scans something? How do you even go about enforcing such a thing?
- brookst 1mo agoWait, so the library of congress is suddenly responsible for probably petabytes a day of incoming scans? To what end? Do they have to index it and make it available? Do they have to check the accuracy and integrity of the scans? How does this help anything, except create more work to throw in the trash?
- ro_sharp 1mo agoThis is already required for new books published in the US, and has been for more than one hundred years. It’s called “mandatory deposit”
- toast0 1mo agoCase law seems to be that mandatory deposit is unconstitutional, fwiw. https://en.wikipedia.org/wiki/Valancourt_Books_v._Garland https://en.wikipedia.org/wiki/Valancourt_Books_v._Garland
- shagie 1mo agoThat revolves around print on demand for books that are out of copyright or where the copyright has been abandoned. > Background Valancourt Books is a print-on-demand independent publishing house specializing in rare and out-of-print books. Valancourt had not registered its books for copyright as the Library of Congress already had original-edition copies of the books Valancourt republishes and any new material in its publications was limited to notes and introductions. If you want a physical copy of The Sorrows of Satan, you can buy it from them. Their argument is that the Library of Congress already has a copy of the book ( https://search.catalog.loc.gov/instances/a0f8fcfe-a255-55d2-a7ac-0889fc7be2e5?option=keyword&query=The%20Sorrows%20of%20Satan https://search.catalog.loc.gov/instances/a0f8fcfe-a255-55d2-... ) and having them deposit it again would be unnecessary. > The Copyright Office has stated that it would modify the language of its deposit demand letters and withdraw its demand for copies if the Copyright Office was notified of the copyright's abandonment. > Several legislative changes have been proposed to address all elements of the case: changes to Section 407 to tie some legal benefit to the deposit, monetary compensation to copyright holders for depositing books, and regulation for a simple and costless method of copyright abandonment. That doesn't change that if you were to publish a book today (or for that matter, have published a book in the past 100 years in the US), you are required to deposit a copy of the book with the Library of Congress.
- shevy-java 1mo ago> From a historical precedent standpoint, this is akin to the burning of the library of Alexandria Let's view it realistically here: AI companies are parasites. Them destroying books to dumb down mankind, absolutely fits into the destruction of the library of Alexandria. Having said that, I think the day of physical hardcopy of books, is not necessarily over, but will be heavily complemented via digital storage. For instance I only keep books that I may re-read later or read many more times, e. g. thick science books. Many other books I can keep as .pdf file without a problem.
- Aurornis 1mo ago> From a historical precedent standpoint, this is akin to the burning of the library of Alexandria, These comparisons are starting to get ridiculous. Why are so many people assuming there is exactly one copy of all of these important books available, that it’s sitting in the warehouse of a bulk book reseller, and that Anthropic is destroying the lone copy? Your local library throws out books every year and nobody thought twice about it.
- snickerbockers 1mo ago>Why are so many people assuming there is exactly one copy of all of these important books available, and that Anthropic is destroying the lone copy? Why are you assuming that each book gets scanned exactly one time and then never again? And why are you assuming that out-of-print books remain easily accessible so long as not every copy has been destroyed? >Your local library throws out books every year and nobody thought twice about it. When they're damaged beyond hope of repair from decades of wear. As for books that haven't fallen apart from being used as intended and merely are no longer desired, my local library generally sells them off at bargain prices. Of course I can't speak to your local library.
- Goronmon 1mo agoAs for books that haven't fallen apart from being used as intended and merely are no longer desired, my local library generally sells them off at bargain prices. What happens when the books don't sell?
- p-e-w 1mo agoThey are thrown away, which has been happening since forever, and nobody ever gave a fuck until AI got involved.
- snickerbockers 1mo agoI think it hits different when you're buying large quantities of used books with an eye towards ones which aren't readily available online, and with the specific intention of destroying them. If all these companies were doing was burning star wars tie-in novels and harry potter sequels nobody would care. That's not their goal because they already have those in their training set. The whole point here is to find rare or underappreciated books from the pre-digital era which nobody ever made publicly available in a digital format. BTW destroying them isn't even necessary for scanning. It's the easiest way because removing the binding and turning it into a flat stack of papers solves many problems but there are actually dedicated book scanners designed to hold open the book while its photographed, and un-curling pages in post-processing was already a solved problem long before people were using AI to correct images.
- thesdev 1mo ago> working towards the benefit of humanity is not an exclusive right / domain of theirs That's not their goal or else they wouldn't be burning books. Their goal is making money no matter the cost to the society.
- bmelton 1mo agoIf they were just chopping them up without scanning them first, then sure, but I think that scanning books and burning books are polar opposites
- logseman 1mo agoBurning books and destroying them in a way that nobody else can access the content anymore is a distinction without a difference.
- jan_m_savage 1mo agoThey have already gotten to Archive.org. Books that were available to borrow are no longer 'available'. SMH
- JKCalhoun 1mo agoI've scraped all the stuff I am interested in. There's a whole r/datahoarders so I'm not alone. ;-)
- himinlomax 1mo agoThey're not destroying rare manuscripts or incunables. They're destroying one (1) copy of a mass-produced item for each AI company. Public libraries destroy millions more yearly as a matter of routine. This is just part of a CCP-aligned moral panic, along with the water use nonsense, and similar with the soviet-aligned moral panic that destroyed the civil nuclear industry 40 years ago.
- TofuLover 1mo ago> This is just part of a CCP-aligned moral panic, along with the water use nonsense What was nonsense about water use?
- brookst 1mo agoIt was never that dramatic, and it’s declining day by day. It’s a panic over a real but small problem. Order of magnitude more water is lost from wasted irrigation (e.g. during rain, of fallow fields, sprayed into windy air, etc) than data centers.
- SXX 1mo agoProblem with data centers is that companies want to build them near densely populated areas that already have problems with water supply and high utility bills.
- himinlomax 1mo ago1. They don't HAVE to use water. Air cooling, closed loop cooling, waste-water cooling, and so on, are options. Easy to regulate. Evaporative cooling is more energy efficient though, but a complete non-issue in places with abundant water and a non-option elsewhere. 2. Datacenters have been shown to reduce utility prices. They provide suppliers with previsible long term demand which allows for cost-effective network and production planning.
- alex43578 1mo agoThe outcry over data centers using a fraction of the water used for things like golf courses or growing alfalfa in a desert.
- brookst 1mo agoNo, that’s the hyperbolic reaction clickbait wants from you. Not all rare books are valuable. Someone’s self-published junk sitting in the garage is NOT analogous to the library of Alexandria. Many, most, maybe all of these “rare” books are being scanned instead of just being recycled. Not a big Reddit fan but there was a great post there from someone in the book industry talking about how non-industry people often give this great moral weight to ever book in a way that is totally disconnected from reality.
- dd8601fn 1mo agoI don’t think any of these people know. All I’ve read, as far as sources go, is a number of rare book sellers saying they’ve had a big uptick in huge orders with no price haggling. Apparently that’s peculiar. And some of them seemed a little concerned. Now I’m certain they’re not chopping up Davincis notebooks, but I’m not certain there aren’t some that would make people wince. And I don’t have any reason to think some reddit librarian knows what’s going on, if anything, either way.
- andsoitis 1mo ago> Now I’m certain they’re not chopping up Davincis notebooks, but I’m not certain there aren’t some that would make people wince. Then they should list the names of these books otherwise I say they're alarmist.
- dd8601fn 1mo agoIf they have, I haven’t seen it. It’s very possible I’ve just missed deeper reporting, obviously. But otherwise I agree. I’m neither losing sleep over it or just trusting that these (historically kinda scummy) businesses are actually behaving. If there’s a serious problem I’d like to see something more concrete. Same for hand-waving the question.
- pessimizer 1mo agoI heard from the first stories that virtually all of these books being ordered have ISBN numbers. Books that are rare that have ISBN numbers are rare because no one wanted them 99.9% of the time. Somebody wants every book, but you'd spend many, many years finding that somebody.
- Leynos 1mo agoIf you tell people who need digital text from books that they need to destroy books after scanning them, they're going to use destructive scanning and destroy the books.
- chrisjj 1mo ago> Despite the copyright restrictions that are forcing companies to do this There are none.
- LaGrange 1mo ago> I was surprised to read that Anthropic (and probably other data / model companies) are doing this and it's extremely disappointing, as working towards the benefit of humanity Look, others talked about how this fetishising of paper books is quite silly (though I don't like it when the destructive scanning is just so one could feed it into a chatbot) but I have to say, all I can do after reading the above sentence is laughing bitterly. Anthropic is an American for-profit company, any talk about "working towards the benefit of humanity" is just marketing lies, and it's always incredible to see people treat those seriously. _Of course_ Anthropic does that.
- michaelsbradley 1mo ago> From a historical precedent standpoint, this is akin to the burning of the library of Alexandria Are you referring to the burning of the Serapeum in AD 391 or the warehouse fires in 48 BC?
- JKCalhoun 1mo agoIs there a difference with regard to the metaphor?
- michaelsbradley 1mo agoDepends on the point being made about “historical precedent” and the lessons to be drawn from such. Also, helps to clarify what exactly the commenter was referring to and possibly help distinguish the centuries-spanning decline of the Library of Alexandria from the violent fate of the Serapeum.
- JKCalhoun 1mo agoMy take was generally: the loss of Alexandria's collection represents a calamitous loss to our collective culture. (But my knowledge of Alexandria extends only to episodes of "COSMOS" and "Connections").
- michaelsbradley 1mo agoIt's a fair point. I honed in on "the burning of" (original comment) versus more generally thinking in terms of "the loss of", because parent context here is "AI companies destroy…".
- deleted 1mo ago[deleted]
- joshstrange 1mo ago> From a historical precedent standpoint, this is akin to the burning of the library of Alexandria, where centuries of knowledge was destroyed and leaving a limited version of the history, the surviving one, and depriving successive generations of significant amount of latent knowledge. How can anyone say this with a straight face. The knowledge is not destroyed, it is transformed. You can make use of it today in the form of LLMs and the scans still exist. Nothing was lost. It's literally no different from them buying books and stocking them in a private library not open to the public. It's not called the Scanning of Alexandria because if it was, it wouldn't have made a blip in the history, Alexandria's libraries were burned, those books, that knowledge was destroyed. Then only thing being destroyed here is physical copy (again for the people in the back: a copy). > Despite the copyright restrictions that are forcing companies to do this, they should maintain archives that are publicly available. Those same copyright restrictions are exactly what would prevent them from sharing the archives. Your beef is with copyright, not the AI companies who are (in this one, rare, instance) following copyright laws/rules.
- pmarreck 1mo ago> From a historical precedent standpoint, this is akin to the burning of the library of Alexandria Hyperbole much? Does the fact that they're being converted to an immutable digital permanent record for all time mean anything to you? Because as far as I know, the works lost to the Library of Alexandria were wiped out of existence, not simply transformed into a more durable form!
- sensanaty 1mo agoThe idea that these corporations or any of the literal sociopaths that work for them give the slightest bit of a shit about "benefitting humanity" is hilariously naive. The one and only thing these entities care about is money, and making as much of it as they can. If they could get away with it, they'd commit every crime that exists if it meant they get a quarter of a percentage increase in their quarterly earning reports.