3 ms·
It is not a big deal. Since the invention of the printing press any important book has been duplicated by thousands, tens of thousands or even million of units
by cladopa 1mo ago
It is not a big deal. Since the invention of the printing press any important
book has been duplicated by thousands, tens of thousands or even million of units.
Just taking one of those and "destroying them"(it is not destroyed, a digital copy with the ability of doing millions of copies is stored somewhere) is not problematic for Humanity.
By the way, I always search for second hand books. Most of the books there are garbage. Most people clean their shelves with the books they don't care about, but preserve the ones that are good. If they are young people that inherited a house and don't care about books, they pick and sell the good ones, giving away the bad books.
If you go to a recycling centre, the garbage to quality ratio is over 100 or more. That is, for every 100 books that are garbage there is one good quality book. It is very rare to find a jewel there.
- brightball 1mo agoWhenever my wife wants to visit antique stores, I always look for old books. I have found several 100+ year old gems.
- pshirshov 1mo agoRead more about this. Depends on the definition of the "big deal" but from what I can understand the problem is that they buy rare things - which exist in just several copies - and they tend to buy _all_ copies.
- dataflow 1mo agoWhere did you see they tend to buy all the copies? This comment is the first time I've heard of this.
- wmeredith 1mo agoI'd also be curious about the provenance of that statement. Why would they buy all copies? What would be the purpose of scanning multiple copies?
- dataflow 1mo agoI could see buying multiple copies being useful to mitigate problems, like damage. But buying all the copies is categorically different and I cannot imagine why they would attempt that, except perhaps to prevent their competition from getting a hold of the same text?
- p-e-w 1mo agoIt’s just another lie of the type these threads tend to be filled with nowadays. Of course they aren’t buying “all copies”, and that wouldn’t even be possible in most cases since such books are usually flea market/attic material and most copies aren’t for sale (or even catalogued) to begin with. I’d be interested to learn who comes up with such lies though. Is it really just random people venting their frustration, or some kind of organized astroturfing operation?
- diseasedyak 1mo agoIt really does seem like an organized operation, given that it's so prevalent and they all seem to be in lockstep with their specious claims.
- wongarsu 1mo agoI wouldn't be surprised to find out that this outrage is fueled by the same actors as the AI data center water outrage. Whoever they are
- vavos 1mo agoI think these type of lies usually come about as a result of a game of broken telephone and things get exaggerated
- sajithdilshan 1mo agowhat is your source?
- mistercow 1mo agoMost old books that are rare and unpreserved are so because their value is marginal, so nobody has bothered to collect and preserve them. But where did you hear that they’re buying “all copies”? And to what end?
- dbspin 1mo agoThis is a classic mistake. We have no way of estimating the future value of a given book. It's perceived current value (a large part of which is simply obscurity) may be low. But it's future value - to historians, ethnographers, to researchers seeking a specific fact or example of language use or a hundred other things - is literally inestimable. To take a crude example in a different medium - new york in the 90s - widely documented right? Yet, if you want to find high definition video of street life in a given burrough on a given day or year, you're faced with an enormously difficult task. There were some HD test videos done in Manhattan in the late 90s (which have been posted to Hackernews before), but there's no equivalent for the other burroughs. Your best best would be finding original negative out takes or location scouting footage from feature films, a very hard task. That's only 30 years ago. Outside of the focal points of the worlds attention - English language, rich countries, places in the news, contemporaneous sources for 'non notable' events (lifestyle, how people spoke dressed etc) is surprisingly poorly preserved. Hopefully you can infer how this tracks to the written world and primary sources for language, technical manuals etc etc.
- famouswaffles 1mo ago>Hopefully you can infer how this tracks to the written world and primary sources for language, technical manuals etc etc. I honestly can't and I think you can't either or you would have used an example with books/printed media rather than film, an entirely different ballgame.
- dbspin 1mo agoOK... I'm going to assume good faith even though your wording makes it somewhat unlikely. Similar textual examples would be any text containing actual language as it's spoken in a given place or time. Or any factual textbook detailing the buildings present in a given location. Or any text book detailing a now defunct construction process. Or any text book (generally small run) detailing a niche interest, now missing ecosystem or the state of a particular political situation at a given time. Essentially all textual primary sources for events which are not currently considered important - but which we have no way of estimating the future importance of. One can continue to create countless counterfactual examples in this vein. My overall point is we cannot know what may be useful or even essential in the future, and knowledge should be preserved under the assumption that it is likely to be. Historians frequently refer to this paradox - how everyday aspects of life are frequently not explicitly documented, since they're so obvious to the communities or communities of expertise that observe and carry them out. So it's actually incredibly important to preserve what seems like ephemera. Hell we couldn't have AI training at all if we lacked the corpus of existing written literature - but there was no way any author could have anticipated this future utility more than a couple of decades ago.
- pfdietz 1mo agoHere we have another entry in the long list of "things described on the Internet that never happened".
- wasmperson 1mo agoI was also skeptical of this claim but managed to find someone who explains it: https://downtownbrown.substack.com/p/five-fallacies-ai-and-destructive-scanning-of-books https://downtownbrown.substack.com/p/five-fallacies-ai-and-d... It's not that individual companies buy all copies of a given book, but that there's more than one book scanning company, and they aren't sharing the scans with each other. The result: books that were rare but nevertheless easy to find for purchase (thanks to the internet) are now vanishing off of the market, becoming de facto no longer accessible to the public.
- quietsegfault 1mo ago[dead]
- demibabs 1mo agoGood article, but I still feel unsatisfied because even it cannot find an example of a book that’s actually been lost because of the destructive scanning frenzy (it only lists books that hypothetically could be lost because there’s not many physical copies available for sale online.). If anyone has an example, I’d love to hear it.
- voidhorse 1mo agoSince we don't know what was actually purchased and what was actually destroyed, how do you expect us to furnish an example? This would require the destroyers to admit it, and beyond that it would require all of them to admit it since more than one of them might have been responsible for the extinction of one text. Seeing as they were already keeping this operation under wraps, I don't see that happening. "possibly extinct because no copies available online" is probably the best we can do. The distributed nature of the problem and the utter lack of transparency are huge factors here too.
- demibabs 1mo agoThese companies are buying from small book stores; surely these collector types would know if they lost any one-of-a-kinds? Or at least someone would be keeping track of extremely rare books disappearing (especially now since this matter has been public for weeks). Regardless I feel like they gotta figure this out for optics reasons. “So and so books are lost forever to Anthropic’s servers” is much more outrageous than “Anthropic is destroying a bunch of books that have other copies” imo.
- shiandow 1mo agoSomehow I don't think they're looking for the books that have been copied over and over.
- quietsegfault 1mo agoWhy do you think that? Do you have evidence, or is this just a hunch? Why would Amazon waste money on uber rare books when there are thousands and thousands of not-so-rare books that could serve the exact same purpose?
- shiandow 1mo agoFor one they've already used the entire library genesis. Anything not in there is going to be obscure in some capacity.
- sajithdilshan 1mo agoExactly, also all those physical books would anyways get molded, eaten by moths or just naturally decay. It's not like the AI companies are obliterating every copy of every single book.
- etdznots 1mo agoWe should actually thank the AI companies for helping us get rid of our trash! Thank you anthropic! Thank you OpenAI! Thank you Google!
- nloomans 1mo ago> a digital copy with the ability of doing millions of copies is stored somewhere somewhere were we can't access it. as the article states: “permanently locking human knowledge inside private corporate servers” the issue isn't that the physical copy is gone, it's that they are preventing people from making digital copies that are actually accessible by destroying the physical copies. > If you go to a recycling centre, the garbage to quality ratio is over 100 or more. archivists keep everything, because we don't know right now what will be important 100 years from now.
- radu_floricica 1mo ago> somewhere were we can't access it By any metric imaginable, it's making the information more accessible, not less. First, it's taking a single copy of a 10k physical print and it's making it digital. Is it "locked"? Yes, by copyright laws, if you don't like that lobby to have them changed. But it's _closer_ to being widely available, not farther. Plus having the info part of a LLM makes it immediately available to literally billions. I happen to actually actively shop second hand bookstores, so I am potentially affected by this - as opposed to most people complaining because they don't like the idea. And I still absolutely support it.
- jjulius 1mo ago>Plus having the info part of a LLM makes it immediately available to literally billions. Help me understand how. Not only are these LLMs expressly prohibited from specifically regurgitating copyright works if the users asks them to, but they habitually hallucinate or paraphrase things wrong. If they won't regurgitate the copyrighted text verbatim, and are known to be confidently incorrect and hallucinatory, I'm struggling to see how these texts are "immediately available to literally billions". And I ask this as someone who has had LLMs give me incorrect assertions about the contents of books.
- SkyBelow 1mo agoDepends upon what you want. For example, knowledge about how the book smells when you open it, something that reader do talk about enough I don't think this should be a strawman, is lost. But, that is about the experience of reading the book, not the knowledge of the book. The exact text? Yeah, I think that is largely lost as well. This is a summary. And for rarer books, it will be a particularly bad summary. The basics of the book are being captured in a space that the right question that retrieve it, but worse than a sparksnote and any well read reader will tell you all the sorts of things a sparknotes already loses compared to reading the book directly. But, that little bit of data is a bit more data than existed before, and future LLMs should get better at giving the information. So, in that sense, the knowledge is better being spread compared to copyright where the book stays in a warehouse until it is disposed of. If it was between this and a sparksnote of the book being made, the sparksnote is far better, but between this and the book simply being disposed of, then the LLM is better but far, far from great. That's a lot of assumptions that goes into the judgment, which is probably why different people reach different conclusions. One person is imagine the alternate fate of this book being slowly rotting in a landfill, the other resting on a bookshelf where it is read at least once fully and then flipped through time to time, and neither are wrong.
- rvz 1mo agoFirst of all, it IS destroyed and it is a big deal. Hardcover copies of books especially 1st - 2nd edition ones (even with mistakes) are rarer than digital scans. Maybe the Bodleian Library at Oxford University should give all their rare books to AI companies to scan and destroy them since it is not a "big deal" anyway. Except that when they did do a pilot with OpenAI to scan these rare books, [0] they did NOT destroy them. I wonder why? [0] https://www.bodleian.ox.ac.uk/services/research-partnerships/digitisation-project https://www.bodleian.ox.ac.uk/services/research-partnerships...
- kccqzy 1mo agoWhy wonder? The answer is abundantly clear if you follow the news. If a book is copyrighted under U.S. law, scanning and destroying counts as a format conversion which qualifies it as fair use, so there is no need to negotiate with copyright holders. See Judge William Alsup’s decision. If Anthropic did not destroy the books after scanning it would have not won the lawsuit, and scanning would be illegal. If a book is already out of copyright then of course they do not have to destroy it afterwards.
- deleted 1mo ago[deleted]
- quietsegfault 1mo agoWhy is it a big deal? Do you think there's no difference between the books curated at Oxford University and the crap that Amazon is buying?
- voidhorse 1mo agoWhy would amazon buy "crap"? Surely they want their model to succeed and they want to train it on valuable input, no? They have more than enough resources to determine whether or not the books are worth buying. They have been in book selling for a long time.
- quietsegfault 1mo ago[flagged]
- afpx 1mo agoI think you may be greatly underestimating the long tail. Several times a year I read sources that reference older books that I can't find online. When I am able to locate them, they often cost at least several hundred dollars, sometimes into the 10s of thousands.
- quietsegfault 1mo agoDo you think that you are somehow special and unique in needing these books? If the books cost in the 10s of thousands, then there's obviously value to other people. I've seen no evidence that Amazon or others are buying $10k books to scan into their corpus. All evidence I've seen is that they're scanning cheap books with no current value and no clear use to people today. I have volunteered with a library, and probably threw hundreds of books over a couple week engagement from a university library into a shredder at the direction of a professional, academic librarian. Libraries are constantly culling books, the EXACT category books we're talking about here (old, never-read). This is happening at a much larger scale, so I would recommend railing against university librarians in addition to the AI juggernauts.
- hughlilly 1mo ago> I've seen no evidence that Amazon or others are buying $10k books to scan into their corpus. Have you seen evidence that they’re buying only readily available books that are plentiful on the market?
- afpx 1mo agoI'm just stating that not everyone is reading the most popular million books. And, there are many, many millions of books in the long tail. A few years ago, after running into this issue several times, I looked into the economics of it to see if there was a opportunity to republish digitally. I found the sales data for some, the ones that had value were selling on average for around $50-150 dollars. Because the typefaces weren't modern, OCR wasn't scalable. Because only 20-40 were transacted each year, it wasn't worth the time. What kind of books are they? In my case mostly historical documents by some relatively unimportant person who was highly important for a very, very niche subject. They still contain valuable, irreplacable information. An analogy: imagine you discover a really cool video game from 20 years ago. You love it and want to find the developer's previous work. But, you find out that the company was purchase by another company which was purchase by another, etc, etc. Sure, maybe the game still exists in some digital vault. But, more than likely it's gone because old things only seem have value these days if some influencer broadcasts it. It's interesting that libraries are purging them. Several times, I've found that the only available copies were at some random university rare book collection in middle of no where, and basically impossible to access unless one fly in - which isn't worth it. They are often donated by a benefactor and stuck there - which is why they're never read. It's not that the content isn't valuable.
- jll29 1mo agoBeware that the notion of "quality" is entirely different for AI companies: they don't seek entertainment, but sentences in a language to train an LLM.
- mannyv 1mo agoI have books, but they are just objects. They're nice objects, but just objects. Fetishizing books isn't going to help. In fact, most of those "rare" books don't sell because nobody wants them. The AI companies are making them even more rare, so the booksellers should be thankful.
- GreenLightGo 1mo agoHonestly, it’s easier to find a good movie than a good book, because books are way cheaper to publish. These days, the quality of pretty much all kinds of content has become a problem...
- TFNA 1mo ago> any important book has been duplicated by thousands, tens of thousands or even million of units. Books in the former USSR display their print runs on the last page. "Important book" is a vague and arbitrary term, but rhere are works in whole fields (e.g. history, archaeology, linguistics, ethography) that any scholar would consider key references, and as few as 100 copies were printed. The shadow libraries have made a lot available to the whole world. It would suck if private corporations scan and shred remaining copies of these before the shadow libraries can get a scan.
- the-mitr 1mo agoOur little project endeavours to digitise books from the erstwhile Soviet state which were published in many languages We have, over the last 15 years, acquired/borrowed from libraries, and scanned a couple of thousand books on all topics of interest. All of them are out of print. Some of the physical copies were have, especially in indicate languages might be some of the few surviving ones https://mirtitles.org/ https://mirtitles.org/ https://archive.org/details/@mir-titles https://archive.org/details/@mir-titles
- pibaker 1mo ago> any important book has been duplicated by thousands, tens of thousands or even million of units. It is common for academic books to have publication runs in the low three digits. You may argue these books are not important. But how do we know if we fail to preserve it?
- quietsegfault 1mo agoIs there evidence that these mythical low-print-run books are being purchased by Amazon and friends for destructive scanning? I simply don't see why railing against the AI giants about this without also contributing significant effort into protecting these low-print-run books from getting jettisoned by university libraries makes any sense.
- etdznots 1mo ago> Is there evidence that these mythical low-print-run books are being purchased by Amazon and friends for destructive scanning? Is there evidence that they aren’t? All of the common works are already on libgen, if these books are worthless repetition they will contribute little to the training corpus, labs want high quality interesting texts and they have unlimited budget to spend on it. Paying $300 for something rare with millions of tokens of interesting and original text for training is definitely fucking worth it, spending $1,200 to get all the copies and block your competitors from getting it is most definitely worth it! > I simply don't see why railing against the AI giants about this without also contributing significant effort into protecting these low-print-run books from getting jettisoned by university libraries makes any sense. Do you think it would be OK to rail against for example, a genocide if I wasn’t substantially contributing to some effort to stop it? This sophistic (and uninteresting) bit of rhetoric boils down to: if youre not trying to fix it yourself, dont complain! (reminder that one of the basic principles of a democratic society is that each person has some concern outside of their personal affairs)
- quietsegfault 1mo agowe're done - straight to genocide, later buddy.
- arttaboi 1mo agoWith all due respect, I would say it wouldn’t hurt not to downplay this.
- XorNot 1mo agoSure but are we going to do anything sensible about it or is it just a convenient culture war vector? This is happening because copyright means you can't scan these without destroying them as a format conversion. No one's felt compelled to try and fix that so we can do this sort of digital archival and preservation, and copyright allows works to be frozen and undistributable because a claim might exist for decades without any actual use (I.e. the number of games which get stuck in legal limbo). If the only desire is to sling mud at AI companies but not try and improve the legal situation, then it's worse then useless because there's no intent to stop it - in fact stopping it would remove a useful outrage tool. The idea in the title here is point 1 of the blind leading the blind: it's illegal to scan and store these books without destroying them in many jurisdictions.
- mtkd 1mo ago>It is not a big deal have you ever held and read an old book?
- NoMoreNicksLeft 1mo ago>Since the invention of the printing press any important book has been duplicated by thousands, tens of thousands or even million of units. No, every popular book has been duplicated thousands of times. This is not the same thing as important. They're orthogonal. When an important book is popular, it is safe. When it is not, it is in danger. Only fools assume that important books are recognized often enough to become popular.
- anigbrowl 1mo agoI don't know why you think all these irrelevant remarks somehow support your core thesis.