4 ms·
Because it provides incentive for people to help store the archive, and it solves the problem in their article about defending against bad actors.
by phacops 12y ago
Because it provides incentive for people to help store the archive, and it solves the problem in their article about defending against bad actors.
- sfeng 12y agoBut all you need to defend against bad actors is for archive.org to hang on to a SHA256 or 512 for every chunk. Much simpler than a distributed blockchain.
- phacops 12y agoIt’s a little more complicated than that. I could temporarily take a chunk, compute the hash, then throw it away and report back that I am happily storing the data, even though I’m not. Anyhow, read the permacoin paper. It’s pretty cool, and it needs a large petabyte data seed to secure the network. Seems like a win-win to me.
- jerf 12y ago"I could temporarily take a chunk, compute the hash, then throw it away and report back that I am happily storing the data, even though I’m not." This is not a new problem, and most of the obvious solutions work just fine.
- roganartu 12y agoThe centralisation helps again here too though. The central server has access to the entire file, and hence can compute the hash of any arbitrary chunk. Challenges/verifications don't have to happen all that often ([1] indicates they are looking at once a month) so creating a unique challenge for each user shouldn't be too compute intensive. For each known user: - Central server chooses random chunk of each file it wishes to verify the user still has. This could be any length from as long as the hash function to the whole file size and could be offset any number of bits. - Client is asked to provide the hash of the chunk using the file stored locally. Precomputing hashes for all possible chunk permutations would take up substantially more space than simply storing the file in the first place. A bad actor would need to store the hash for all the possible chunk lengths starting from every possible start location in the file which is in the order of O(n^2) where n is the stored file size (500GB in this case). For reference that would be about one third of the entire 20PB archive for a single 500GB chunk if using a 256 bit hash function. [1] http://git-annex.branchable.com/design/iabackup/ http://git-annex.branchable.com/design/iabackup/
- darkmighty 12y agoIt's a good thing in this case they also don't have to worry too much about bad actors, since it's a fundamentally altruistic endeavor. Worst case scenarios: 1) Someone keeps downloading the full archive and throwing away; 2) Someone wants a file erased, keeping it with the intention of denying access in some future. -- 1) Bandwidth has costs on both sides; just balancing upload among receivers would probably suffice for this not to be a problem. 2) Assigning large random chunks to downloaders should prevent this "chosen-block attack"; add in some global redundancy for good measure and that's probably enough (although I still wouldn't trust this 100% as a primary storage, only as an insurance storage).
- shabble 12y agoThe added bonus of periodically validating random chunks is that it forms a sort of data-scrubbing function, allowing the detection of non-malicious failures due to bad sectors or whatever as well.