4 ms·
Jason Scott here. Just wanted to address the questions that always come up when this project gets some attention. (Also: Come volunteer to be a client! The more
by textfiles 11y ago
Jason Scott here. Just wanted to address the questions that always come up when this project gets some attention. (Also: Come volunteer to be a client! The more the merrier.)
* We are only backing up public facing data. (Roughly 12pb)
* We are only backing up curated sets of data. (So less than that.)
* We are stepping carefully to learn more about the whole process as we go, documenting, etc.
* The hope is this will produce some real-world lessons and code that other sites can use.
* This project uses non Internet Archive infrastructure, and is not an Internet Archive project.
It's going well, and the more people who join up, the better. Oh, and support the Internet Archive with a donation - it's a meaningful non-profit making a real difference in the world. http://archive.org/donate http://archive.org/donate
- meesterdude 11y agoAre you guys concerned with decay at all? How do you know the file stored 2 years ago is still the same file? what happens if its become corrupted?
- textfiles 11y agoCurrently, the system requires you to check in (using the git-annex facility) on a regular basis. If you don't check in (it's a script run by cron) within 2 weeks, you're considered decayed, and after 30 days, your contributions fall back into the pool and are taken in elsewhere.
- meesterdude 11y agoI'm talking more about the checksums of files over time, and error correction should they deviate. I recently read about facebooks cold storage and it got me thinking about it (https://code.facebook.com/posts/1433093613662262/-under-the-hood-facebook-s-cold-storage-system-/ https://code.facebook.com/posts/1433093613662262/-under-the-...) though for them they just want one copy total.
- textfiles 11y agogit-annex is the engine behind this project currently, and a lot of bugfixes, feature-adds and work has happened as a result. https://git-annex.branchable.com/ https://git-annex.branchable.com/
- j_s 11y agoI would be interested to hear about protections against active misinformation attacks, where someone is attempting to change archived content maliciously.
- joeyh 11y agoThe checksums we have from the IA are md5sums, which are not ideal, so a preimage attack is possible, but AFAICS you could only use it if you're generating the original file that is stored in the IA, and are planning to replace that with a colliding version in the future. Otherwise, we can detect falsified files. There are some potential attacks of putting false information into the git-annex repositories, that we use for tracking which clients are storing which files. We'll eventually need post-receive hooks to validate that pushes only change information about the client making the push. For now, if someone attempts this attack, we can revert their malicious changes after the fact.
- gojomo 11y ago...if you're generating the original file... Maybe for the time being, with published attacks. But MD5's strength against attacks has been sufficiently dubious since early 1996 – before the founding of the Internet Archive – that experts have recommended against using it in new applications. By 2009, the CMU Software Engineering Institute (responsible for CERT security notices) wrote: Software developers, Certification Authorities, website owners, and users should avoid using the MD5 algorithm in any capacity. As previous research has demonstrated, it should be considered cryptographically broken and unsuitable for further use. So they're worse than 'not ideal'. They're dangerously obsolete for the purpose of content-authenticity, and any reliance upon them (under hand-wavy conditions) serves to block the migration to available hashes that would actually work. If you're making the backup to be sent through a wormhole to 1999, use MD5. If you're creating the backup for trust and availability through to end-of-year 2015, or 2025, or 2115, ditch MD5 ASAP. [1] http://www.kb.cert.org/vuls/id/836068 http://www.kb.cert.org/vuls/id/836068
- nrao123 11y agoThanks for doing this. Hopefully, this will help reduce the problem of the Web of Alexandria that Brett Victor talked about: 60% of my fav links from 10 yrs ago are 404. I wonder if Library of Congress expects 60% of their collection to go up in smoke every decade. --- For someone who's thinking about a library in every desk, going on the web today might feel like visiting the Library of Alexandria. Things didn't work out so well with the Library of Alexandria. It's interesting that life itself chose Bush's approach. Every cell of every organism has a full copy of the genome. That works pretty well -- DNA gets damaged, cells die, organisms die, the genome lives on. It's been working pretty well for about 4 billion years. We, as a species, are currently putting together a universal repository of knowledge and ideas, unprecedented in scope and scale. Which information-handling technology should we model it on? The one that's worked for 4 billion years and is responsible for our existence? Or the one that's led to the greatest intellectual tragedies in history? https://twitter.com/worrydream/status/478087637031325697 https://twitter.com/worrydream/status/478087637031325697 http://worrydream.com/TheWebOfAlexandria/ http://worrydream.com/TheWebOfAlexandria/
- thaumasiotes 11y ago> It's interesting that life itself chose Bush's approach. Every cell of every organism has a full copy of the genome. Ehhh... every cell has a full copy, but that's more of a coincidence than anything. They're not capable of using their full copy. And only the germ cells make any contribution to the genome of the next generation.
- ISL 11y agoEvery cell has a full copy of the current operating plan, not the entire history of all preceding operating plans. Storing the entire commit history of our DNA would be much more space intensive.
- ddlatham 11y agoHow much more?
- 11y ago