4 ms·
Quite possibly not many, and no copy held in any form by the copyright owner either. Say a few hundred copies of some obscure book from 40 years ago. They proba
by card_zero 2mo ago
Quite possibly not many, and no copy held in any form by the copyright owner either. Say a few hundred copies of some obscure book from 40 years ago. They probably won't be erased from the face of the earth by the judicious and proportionate actions of, of a few, AI companies? Hmm.
- willy_k 2mo agoIs there a specific book from 40 years ago you have in mind? Asking out of curiosity.
- ipaddr 2mo agoBooks by Zolar are interesting hard to find all editions. The Fearful Void by Geoffrey Moorhouse probably still has 100s of copies available but hate to see it lost.
- scarmig 2mo agoThe hypothetical "heroic figure goes and buys last copy of a 1962 guide to Ford cars to carefully maintain it in an appropriately climate controlled library" is vanishingly unlikely. A ten or a hundred or a thousand times to one, it just goes to the trash. At least here it gets scanned by the AI company.
- asaddhamani 2mo agoBut that scan is never made available to us in its original form. So it getting scanned by the AI company does nothing to preserve the book.
- scarmig 2mo agoDumpsters also don't typically come equipped with a robot scanner and network uplink built in. Like, I really don't know what people objecting to this imagine typically happens to old, unwanted books. They don't get sent to some magical library in the countryside if unpurchased where they are carefully maintained forever (next to where Rover spends the rest of his days). They are very literally thrown into the trash. That said, I'd be thrilled if the US government required AI companies to make them available to the public. I'd even settle for the US government making it legal for them to.
- soperj 2mo agoThey're buying the books from resellers, not rescuing these books from dumpsters. Stop being an apologist.
- scarmig 2mo agoThe magical library in the countryside, to be painfully explicit, does not exist.
- Ekaros 2mo agoAnd they should not even be needed. In many places the issue is solved at start. Copy or copies of each commercially produced book is send to national library. Which with tax payer money keeps an archive. Meaning that at least one copy exist for research purposes if needed.
- scarmig 2mo agoUnfortunately, that's not the case in the United States. The LOC only selects around half of published books to be permanently held. The rest are disposed of (usually returning them to the publisher, donating them to a library, or destroying them).
- fmajid 2mo agoThey should send them to The Internet Archive instead.
- FeloniousHam 2mo agoWhy aren't we storming the Library of Congress? They are the real villains here.
- fmajid 2mo agoNo, Congress itself is, for not giving its own Library the resources to do is job.
- red75prime 2mo ago...because it is illegal to copy copyrighted material. 70 years later they might do it.
- rhdunn 2mo ago95 years after publication. Many other countries also have an X years after the author's death clause where X varies between countries but is at least 70. There are also other weird issues such as the UK having a clause protecting Peter Pan (so a children's hospital gets royalties) and the King James translation of the bible (under Crown copyright) that extend the copyright even further. In short, it's a mess.
- mrweasel 2mo agoThe thing I find most hypocritical though is that they are probably never share their libraries with anyone. After scanning, downloading, stealing, overloading websites and everything in between, to acquire enough data for their stupid machine, they're not going to share their data? I get that most of it can't be shared, but a lot can. There's no reason why you need to destroy multiple copies of a book from 1880, when it's free to share. At the same time I can understand keeping track of when each books enters public domain might also be an absolute nightmare, and I wouldn't blame the AI companies for not wanting to deal with that. For the stuff they absolutely know is clear, they should provide dumps for everyone to download.
- fmajid 2mo agoSince they seem to leapfrog each others’ models every few months, the training data is one of the few ways they can build competitive advantage, and that explains why they don’t share, even if we don’t have to like this.
- svachalek 2mo agoWould it even be legal to share? I don't think it would be.
- novok 2mo ago
- skeledrew 2mo agoIt was never available to you/us in the original form either.
- dukeyukey 2mo agoIf it were legal they may well do that as a public branding exercise. Google already tried and got punished for it!
- mejutoco 2mo agoIn my opinion this is one of the reasons why libraries should accept any book, even if all they do is examine it and throw it in the trash. This way they would have a chance at finding any treasures that could be regularly dumped in that way.
- bulbar 2mo ago> They probably won't be erased from the face of the earth by the judicious and proportionate actions of, of a few, AI companies? I don't see why not. Pretty sure it's gonna happen. Doesn't matter if a hundred copies still exist somewhere, if access or discoverbility falls below a certain threshold, it doesn't matter, because those books become practically inaccessible to the world.
- margalabargala 2mo agoRight, but if an AI company buys some vanishingly uncommon book, digitizes it, shreds it, and adds the information it contains to their permanent digital library and digests its contents into an AI that is then publicly accessible...are they making that book less accessible, or more?
- scarmig 2mo agoYou've got to compare it to the alternative. Books have a half-life, and the vast majority of these books being purchased are grody, moldering ex-lib copies of books that no one has read in decades. Their other likely outcome is mulching.
- margalabargala 2mo agoRight, that's my point. These generally are not books people care about. The information contained therein was doomed. Now the information has been digitally preserved and a digestion of the information will be made publicly available.
- michaelmrose 2mo ago[dead]
- halsafar 2mo agoCan you get the exact text back out with a prompt or not? Having or not having a book isn't fuzzy.
- qingcharles 2mo agoMany are just copyright "orphans", nobody knows who owns the copyright any longer. Maybe the author died and the copyright passed to their estate, but they're not even aware of it. One book I'm hunting for a copy of right now was published in England in 1947 and in those days paper was rationed, so not many copies were made, and only a handful have survived. As soon as I find it I'll scan it and upload it to IA.