7 ms·
The critical thing here is that regular people can help. The archive is already split into many O(GB)-sized chunks. We need to ensure each chunk has many, many
by pradn 2y ago
The critical thing here is that regular people can help. The archive is already split into many O(GB)-sized chunks. We need to ensure each chunk has many, many copies.
Torrents are a widely-understood, robust way to mirror large files. But there's no "meta-coordination" built in to the protocol. It's not possible, using just torrents, to have the swarm cooperatively assign who stores which chunks. The optimization function here is maximal chunk availability, subject to individual storage limits and reliability (ie: how often they're online).
It should be easy to just press a button to join a shadow library, allocated 100GB, and be part of the mission.
The best effort I've seen in this space is some guy running a script that crawls the number of seeders for a list of SciHub torrents. Users manually pick the ones with the lowest seeds. All very cumbersome, and prone to staleness.
Of course, this is all a technical problem, separate from infringing copyright or whatever. In the same way as torrents being a technical solution for sharing files, in a general way.
- squigz 2y agoThis does seem like it would be the way to handle such libraries, considering the immense size of them. I'd be curious to hear about any efforts in this area, of anyone knows of any.
- throawayonthe 2y ago[dead]
- ric2b 2y agoIPFS is a network that solves your coordination problem, compared to torrents it allows you to decide which chunks you want to store and it will even de-duplicate automatically across different "torrents" that happen to include the exact same byte-identical file.
- pshirshov 2y agoThough with its persistent brainsplits it's barely usable, unfortunately.
- j_maffe 2y agoCould you elaborate? IPFS is already being used quite successfully by LibGen and Z-lib AFAIK.
- pshirshov 2y agoWell, it might take literally months for a file to become available/resolvable everywhere. Also if a host can resolve a file now doesn't matter it would be in ten minutes (the publisher is always running of course).
- pradn 2y agoI like IPFS, and want it to continue improving, but it is just too slow, uses too many resources, and is often unreliable. https://annas-archive.org/blog/putting-5,998,794-books-on-ipfs.html https://annas-archive.org/blog/putting-5,998,794-books-on-ip...
- bhaney 2y ago> regular people can help Can we? I have tens of terabytes of unused storage space that I would be happy to contribute to library archival, but my understanding is that if I seed these torrents I'm going to get spammed with DMCA letters from publishers until my ISP gives up and cuts my service. If we need to seed exclusively through tor or VPNs in copyright-notice-ignoring countries, then that's not all that accessible to "regular people" anymore.
- zozbot234 2y agoI wonder just how much of this "shadow library" content is stuff that's actually in the public domain and could be mirrored with no legal jeopardy whatsoever. Unfortunately, the low quality of catalog-like metadata that's available from so-called "shadow libraries" (including the one that's linked in the top comment) makes this a very hard question to answer. Even if that makes up only a handful of TB's or so, it would be worth mirroring the stuff - among other things, a reliable repository of copyright-free content would also be a valuable resource for "ethical" AI training and other such uses. (I know that the linked blogpost mentions that they just don't bother mirroring "widely available collections" of public domain books. The interesting question is whether there's some "long tail" of PD content that might not have made it to the more well-known collections as of yet.)
- squigz 2y agoFWIW, I've been getting (automated) DMCA notices for years (since 2015 or so) with no warnings or anything like that from my ISP. They just forward the notices because they have to. Not to say there's no risks here, but this really depends on your jurisdiction, I think.
- op00to 2y agoMy ISP, Verizon, has threatened me with shutting off my service after a single notice when my kid discovered bit torrent.
- ptek 2y agoIs remember back in the day 2008 you could download individual files from torrents (Amiga disk images). Is there a torrent toll that can compare individual files in a torrent and if the hashes are correct, download from a bunch and rebuild the torrent. Some torrents add another ASCII advert like Amiga BBS's used to add back in the old day which used to result in dupes?. Amazing that book piracy is a thing, I guess it's pretty big on this board as the majority of users on this board would be considered 'bookish' and know that a lot of the (technical) books that would like are not available at the local library.
- pradn 2y agoYou can still download specific files from torrents of many files, yes. You can pick a subset of files to download. In theory, you could extend torrent clients to support archiving some subset of the files. You just tell it to keep around 100 GB of files in a 10 TB torrent. Since the client knows the status of each chunk in the swarm, it can make smart decisions about which chunks to download and seed. Clients already let to set preferences to seed the "least-seeded" chunks. So this isn't so far fetched. The big problem with this is that it lets you work only with a fixed archive. Torrent files can't be mutated after they are created. So you'd be able to get people to archive a fixed version of a shadow library, but not an evolving one. In practice, this should be fine enough. If we had high assurance that all the books added to the library before 2022 (or something) were copied in 50 machines, that's quite useful. Being able to layer deltas on top would be amazing - you could evolve the collection as new books are added.
- everforward 2y ago> It's not possible, using just torrents, to have the swarm cooperatively assign who stores which chunks. I don't know if it's strictly desirable to have the swarm cooperatively assign who stores which chunks, because that provides an avenue for bad actors to attack that assignment (e.g. by e.g. claiming to have a block to drive peers away, but never actually serving it). It would be possible to have the swarm behave cooperatively based on heuristics, though. Your client gets a copy of what peers have what chunks, so it has enough information to make its own decision on what chunks need to be mirrored the most. A sufficiently clever algorithm would get pretty close to a centralized cooperation server. Iirc, some extant torrent clients have similar features where they download the "hot" chunks first (the chunks with the most leechers, for private tracker ratios). I suspect the only reason an "archive/sparse" variant of that where it only downloads poorly mirroed chunks doesn't exist is because it's useless for the normal "download a file" use case. Sparse chunks of a file are basically useless outside of archival.
- pradn 2y agoYou’re talking about possible sibyl attacks. Yes that’s a concern but there’s a large literature on how to reduce their threat. This might be a place where “proof of x” concepts from crypto land work. But we don’t need to go all the way to the maximal solution for it to be practical.