3 ms·
Ask HN: Is decentralized historical data preservation possible?
People and states are known for burning books and doctoring records. But those are obviously essential for conducting any kind of historical research, whether in academic, judicial, or other contexts. I am not even referring to secret government communications but rather to public and seemingly banal stuff like old newspapers.
The hope with digitization was that it would preserve all of these artifacts. However, AI approaches the point where imagery and textual data could be doctored on a large scale for pennies, and if data is stored only in a few places controlled by the same entities, altering it becomes trivial. If now I can obtain a scan of an old newspaper from a state library website and trust that it is authentic (because who has the time to professionally photoshop it?), soon there will be no way to be sure since one ill intended person with access could mix things up.
Decentralization seems like a modern solution to this problem, but I wonder if it is sustainable. To run some numbers, if we take the US as an example and just consider the Library of Congress, we would need to distribute 21 petabytes between 2,000-3,000 people with 10TB drives. Therefore, something like this seems to be feasible with a couple of thousand volunteers and $1-2 million for hard drives to start ($50-100 per TB), and then $100-200k per year (assuming a 10-year lifespan of a hard drive). But that's only one archive and quite a lot of people to organize.
To me, it seems that something like this is essential to preserve some form of trust in historical truth. And now is the time to act because otherwise, soon there will always be a second thought that what we are looking at is not real.
- Is there a smart way to solve this problem more easily and cheaply?
- Or am I crazy and we should not be worried about it?
- theGeatZhopa 3y agoMicrofilms. Digital is nice. But what do you do, if there is a big cyber war going on. Microfilms. May be find a solution to make them even more densier and take even more data. The information saved on them can be anything from img/text/encoded whatever. Of course, it's a old technology. But, this can be copied, distributed and decentralized. Just the visual file format needs to be invented haha The archives and libs storing backups of our past are usually well planned by librarians and other guys of topic. I dont think it's really necessary to think about this problems :) the other thing is: What is worth to be archived? What will happen with the massively generated content? How to decide? I see a poisoning of archives in the future:)
- kaurov 3y agoThe problem with physical solutions like microfilms is accessibility. I will be happy to know that there is a “true” copy buried somewhere in a mountain, but I won’t be able to access it. I had no intention to throw any shade on librarians :). They are cool, and do a lot of useful work assembling stuff, but I as user what to have trust in the original source. If everything including backups is managed by one organization, I can be still suspicious. :) And your point regarding what do we actually want to preserve is very valid. It is easy to imagine a situation when we want to store too much.
- theGeatZhopa 3y agoYes. That's a good question... Right now, in U.S. each president's fart must be archived, like the tweets of D.Trump, which has been so massive in counts, that they've had to buy a new harddisk for it :) I'm not sure about America's rules. In Germany, the law for federal archive says "movies, songs, events (like who got the award, who was nominee..), written texts and - actually - everything which is not Facebook or a comment in Instagram. Basically everything that is published and publicly available/accessible. Of course the whole state running thibgs, the prosecution of crimes and the ruling of judges.. Quite a lot of stuff, if you ask me :) but ok! For the archive's sakes! While I'm with you at the most of your arguments, with the digital archive you'll have more problems with accessibility. You'll need a computer for it. The File-Format/codec/code should also be readable by the future's archive software.. (it's actually a problem of itself. One of the formats suitable and already used for documents is PDFX A..) So, while a digital archive can offer you error corrections in case data gets corrupted because of, f.e. medium's degradation, you still are in need to make sure it's copied early and often to other disc, so not too much data is lost for the error correction to be able to work.. With analog, the medium is the problem. A harddisk is typically suitable for 20-30 years, USB Stick / SSD are somewhere in the 2-4 years (if not being used once in a while).. microfilms are 150y+, tone or ceramic plates 1000y+++. In analog one has to take care of the medium itself (restoration f.e.) But with analog you don't need a computer. You need a light source and a magnifier. This will be available after the doomsday too - what can't be said about computer/infrastructure/powerplants.. Actually a good topic for a good night read. Thank you for raising :)
- nonrandomstring 3y agoFor a dynamic solution massive redundancy and some kind of stakeholder accounting. Some kind of "planetary file system" established in the public interest. Are the Wikipedia and Internet Archive good models? Technically a kind of thing like a benevolent version of Web3.0 might solve if it were not motivated solely by profit. IPFS and variants are the landscape at the moment. The consideration (reward) for participation is access to the whole corpus. You need to solve abuse problems which mean Wikipedian-like entry guards and curating of content insertion. Do you trust them? We'd need at least 10x redundancy, so imagine more like a few million people with 10TB to spare. That's not unreasonable in the next few years. But that needs to be maintained or else swathes of data will be lost forever if not enough copies are kept. That means you need at least a few million people who seriously give a shit, and those are getting harder to find each day. Also, with advances in compression, that could be a lot less space if you're prepared to trade-off access time. Like going to the library to find and scan a newspaper, image making a request that takes the distributed system several minutes or hours to lookup, assemble, collect and decode all the pieces. That doesn't sit well with a "give it me now!" culture. For static a solution, burying a few petabytes of long-range storage at strategic (maybe secret) locations is an assurance against Nineteen Eighty Four/Fahrenheit 451 scenarios - a library to outlive the next tin-pot "Thousand year Reich". But that doesn't give anyone quick/random access to disputed facts and records. A more political/human solution is to get more people caring about history and truth, ad-hoc curation and preservation. That requites not just liberation of data a la Aaron Swartz and Alexandra Elbakyan, but hugely increasing the number of people who will participate in that project. At this point, belief in the preservation of historical human knowledge means fighting the law,
- kaurov 3y agoNeither wikipedia or Internet Archive are fully protected from someone tampering their databases. And while I want to believe that they have security sorted out, it demands me to blindly trust them. I like and use Wikipedia but I do not trust it 100% (not because of possible tampering but simply because of the way it is assembled). Giving access to the corpus as a reward is great but the problem is that very few people will actually want to use it :). On the other hand slow access time should not be the problem for those few who care. Maybe the best combination is to continue using fast access systems (like existing digital libraries) and then slow systems only for confirmation. Curation wiki-style might not even be the main issue at the beginning because a lot of things are already pre-curated. Like old radio recordings or newspaper archives. Curation will become a nightmare with modern media like twits… Changing people’s minds would be great, but I am afraid AI is evolving on a much shorter timescale. :)