9 ms·
This seems to get brought at least once in the comments for every one of these articles that pops up. The IA has tried distributing their stores, but nowhere n
by jdiff 2y ago
This seems to get brought at least once in the comments for every one of these articles that pops up.
The IA has tried distributing their stores, but nowhere near enough people actually put their storage where their mouths are.
- immibis 2y agoKeep in mind the IA archives a lot of garbage. If it could be more focused it would be more likely to work.
- db48x 2y agoThe attempts have actually been focused on specific types of content, such as historical videos.
- Blackthorn 2y agoThe IA only works because it archives everything. You don't know what you need until you need it.
- Spooky23 2y agoArchives generally purposefully don’t have a strong editorial streak. My trash is your treasure.
- immibis 2y agoThey have to if they don't want to use infinite space.
- unleaded 2y agopersonally I love all the random crap on IA!
- WarOnPrivacy 2y ago> nowhere near enough people actually put their storage where their mouths are. Typically because most people who have the upload, don't know that they can. And if they come to the notion on their own, they won't know how. If they put the notion to a search engine, the keywords they come up with probably don't return the needed ELI5 page. As in: How do I [?] for the Internet Archive?, most folks won't know what [?] needs to be.
- TZubiri 2y agoThis is literally torrents. Just give up
- briandear 2y agoThe problem with torrents is they have a bad reputation since people use it to steal and redistribute other people’s content without their consent.
- card_zero 2y agoIs there any form of torrent where you can do a full text search? That, to me, is the more important problem with torrents.
- TZubiri 2y agoBut internet archive doesn't do this? It's a key based search (url keys)
- card_zero 2y agoInternet archive allows full text search of books, newspapers, etc.. Or anyway it did, before being breached.
- TZubiri 2y agoIt does transcribe books (through imperfect OCR) so I guess that's possible. Never relied on it as I search by title and author. But anyways not the case for the wayback product which is the unique core to IA.
- card_zero 2y agoThat's not unique, not a product, and not the part I use most. Well, OK, maybe other webpage archives don't work as well, I haven't tried them, but there are others. And they're newer, so don't have such extensive historical pages. Large numbers of Wikipedia references (which relied on IA to prevent link rot) must be completely broken now.
- creer 2y agoAnd it's guaranteed not to happen if the efforts don't continue.
- acdha 2y agoYou could say the same thing about perpetual motion. Being realistic about why past efforts have failed is key to doing better in the future: for example, people won’t mirror content which could get them in trouble and most people want to feel some kind of benefit or thanks. People should be thinking about how to change dynamics like those rather than burning out volunteers trying more ideas which don’t change the underlying game.
- creer 2y agoThere are certainly research questions and cost questions and practicality and subsetting and whatnot. Addressed by some ideas and not by others. What there isn't is a currently maintained and advertised client and plan. That I can find. Clunky or not, incomplete or not. There are other systems that have a rough plan for duplication and local copy and backup. You can easily contribute to them, run them, or make local copies. But not IA. (I mean you can try and cook up your own duplication method. And you can use a personal solution to mirror locally everything you visit and such.) No duplication or backup client or plan. No sister mirrored institution that you might fund. Nothing.
- zelphirkalt 2y agoPerhaps one idea is to let people choose what they want to protect. This way people wanting to support it can have their mission.
- card_zero 2y agoI want it to protect all sorts of random obscure documents, mostly kind of crappy, that I can't predict in advance, so I can pursue my hobby of answering random obscure questions. For instance: * What is a "bird famine", and did one happen in 1880? * Did any astrologer ever claim that the constellations "remember" the areas of the sky, and hence zodiac signs, that they belonged to in ancient times before precession shifted them around? * Who first said "psychology is pulling habits out of rats", and in what context? (That one's on Wikiquote now, but only because I put it there after research on IA.) Or consider the recently rediscovered Bram Stoker short story. That was found in an actual library, but only because the library kept copies of old Irish newspapers instead of lining cupboards with them. The necessary documents to answer highly specific questions are very boring, and nobody has any reason to like them.
- oxygen_crisis 2y agoYou could let users choose what to mirror, and one of those choices could be a big bucket of all the least available stuff, for pure preservationists who don't want to focus on particular segments of the data. Sort of like the bittorrent algorithm that favors retrieving and sharing the least-available chunks if you haven't assigned any priority to certain parts.
- deafpolygon 2y agoMy favorite question is: whether or not Bowser took the princess to another castle.
- card_zero 2y agoSince the IA had a collection of emulators (some of them running online*), and old ROMs and floppies and such, it could probably help with that one too. * Strictly speaking, running in-browser, but that sounded like "Bowser" so I wrote online instead.
- jonny_eh 2y agoNearly every entry in the library has a torrent file (which is a distributed storage system), but with the index pages down, they're not accessible.
- HappMacDonald 2y agoThey're not using DHT?
- jdiff 2y agoThey're not talking about peer discovery, they're talking about .torrent file discovery.
- highwaylights 2y agoYou're correct, but even then you've still the problem of storage - the torrents are only useful (and there's a lot of them) if a sustainable number of seeds remain available.
- Wheatman 2y agoHow abiut torrenting a collection of websites in one collection? You can distribute less popular websites with more used ones to avoid losing it? And Torrents are good with transfering large files in my experience.
- baby_souffle 2y ago> You can distribute less popular websites with more used ones to avoid losing it? So long as this distributed protocol has the concept of individual files, there _will_ be clients out there that allow the user to select `popular-site.archive.tar.gz` and not `less-popular.tar.gz` for download. And what one person doesn't download... they can't seed back. Distributed stuff is really good for low cost, high scale distribution of in-demand content. It's _terrible_ for long term reliability/availability, though.
- rolandog 2y agoPerhaps a naïve question, but hasn't this problem been solved by the FreeNet Project (now HyphaNet) [0]? (the re-write — current FreeNet — was previously called Locutus, IIRC [1]). Side note: As an outsider, and someone who hasn't tried either version of FreeNet in more than almost 2 decades, was this kind of a schism like the Python 2 vs. Python 3 kerfuffle? Is there more to it? [0]: https://www.hyphanet.org/ https://www.hyphanet.org/ [1]: https://freenet.org/ https://freenet.org/
- sanity 2y agoHi, Freenet's FAQ explains the renaming/rebranding here: [1] Neither version of Freenet is designed for long-term archiving of large amounts of data so it probably isn't ideally suited to replacing archive.org, but we are planning to build decentralized alternatives to services like wikipedia on top of Freenet. [1] https://freenet.org/faq/#why-was-freenet-rearchitected-and-rebranded https://freenet.org/faq/#why-was-freenet-rearchitected-and-r...
- rolandog 2y agoThanks for pointing it out and for correcting me!