11 ms·
The GitHub Arctic Code Vault
- sixhobbits 6y agoWhat stops this stored data degrading? Do they have to periodically check / renewal the reels?
- makerofspoons 6y agoI was hoping for more of a description on how they plan to keep this vault safer than the Global Seed Vault, which was once flooded due to soaring arctic temperatures: https://www.theguardian.com/environment/2017/may/19/arctic-stronghold-of-worlds-seeds-flooded-after-permafrost-melts https://www.theguardian.com/environment/2017/may/19/arctic-s...
- erikbye 6y agoThat was sensationalism, per usual. Bit of water in the access tunnel, no seed damage.
- price 6y agoThe story says right up front in the subhed that the flooding didn't reach the seeds. But the quotes make it pretty clear that what did happen was out of spec. For something that's meant to survive any catastrophe that might happen over centuries to come, it's not a good sign to see that happen so early. It's extra bad to see it driven by a trend, namely global warming, that we're continuing to push farther and farther and have shown few signs of stopping.
- toomuchtodo 6y ago> The Internet Archive is a well-known, widely beloved non-profit digital library which provides free public access to collections of digitized materials. In partnership with the GitHub Archive Program, the Internet Archive (IA) commenced its ongoing archive of GitHub public repositories on April 13 of this year. At present, IA is using a two-pronged approach. First, their well-known Wayback Machine is accessing and archiving raw GitHub data as WARCs, or Web ARChive files. As of this writing they have archived some 55TB of data. Second, they have the goal of making entire archived GitHub repositories available via “git clone,” while also keeping repo comments, issues, and other metadata easily accessible on the web. This second initiative is well underway and initial archiving is expected to commence this month. Tremendous news.
- Gollapalli 6y agoBeautiful. Honestly, nothing scares me more than losing all the code and all the technology we've developed in the past 70 or so years. There's been so much advancement, but it's also transferred in such a way (institutional knowledge, propietary software, proprietary hardware, etc.) that it's super easy to lose. If we preserve open hardware and software, then we could rebuild in the case of civilizational decline and the accompanying knowledge loss, something which we would neither be the first nor the last to experience.
- Wowfunhappy 6y ago> If we preserve open hardware and software, then we could rebuild in the case of civilizational decline and the accompanying knowledge loss ...can we? I'm sometimes a little concerned about how complicated chip fabs are. They feel like something that could take generations to rebuild, even if we had all the knowledge on what to do.
- helldritch 6y agoHome photo-lithography and chemical etching setups aren't common, but have been done by several people. We wouldn't be able to jump straight to 14NM, but we would probably be able to get to the 500-300nm size relatively quickly (a year or two, maybe, if starting from scratch) and shrink down from there. Devices would be much bigger and less efficient, but we would be able to run code and pump out 8086 processors within 6 months.
- quicklime 6y agoThat's just one layer of the stack though. Future archaeologists will also need to create mock npm registries and maven repositories, and set up docker and k8s so they can deploy a complex set of microservices to look up our birthdays.
- Wowfunhappy 6y ago...all the code to which should be right in the Github Vault, right? Idk, the hardware part seems much more difficult to me.
- brendanmc6 6y agoI'm curious, do they perform some sort of test reads on the reels to make sure that the data was actually copied over correctly?
- gdsdfe 6y agoAm I the only one thinking this is a waste of money and time?! How any of this makes sense, maybe as a weird PR stunt but ... Just strange
- dakiol 6y agoI agree. I can't believe they are spending so much money and effort to preserve code I don't give a damn now and once I pushed to GitHub. And like me, 99% of the devs I know personally.
- cmrx64 6y agoit's probably less effort to just archive the whole damn thing and let the future figure it out than to decide important things to archive and leaving everything else to disappear someday
- awb 6y agoI wonder how much space you'd save if you excluded repos with only 1 star or only 1 commit.
- saagarjha 6y agoThey’ve excluded pretty much everything below a hundred stars, from what I see.
- ralph84 6y agoThe inclusion criteria[0] were: > The snapshot will include every repo with any commits between the announcement at GitHub Universe on November 13th and 02/02/2020, every repo with at least 1 star and any commits from the year before the snapshot (02/03/2019 - 02/02/2020), and every repo with at least 250 stars. [0] https://archiveprogram.github.com https://archiveprogram.github.com
- rezendi 6y ago
- chickenpotpie 6y agoSo, if I'm in the EU can I GDPR my repo out of their vault?
- adrianpike 6y agoI'm guessing you're being facetious, but it has come up and it's covered in the FAQ: https://archiveprogram.github.com/faq/ https://archiveprogram.github.com/faq/
- therealmarv 6y ago"... archives have a special legal status under GDPR which protects them. GitHub’s Legal Team has approved the Archive Program."
- therealmarv 6y agoOnly disappointed that the new badge does not show the 2 open source projects I contributed to in the last 10 years of my work for open source :( They are not super big, but also not super small. Seems organisation work is ignored and only individual username fork/PRs respected (is this a bug?). Software is teamwork ;) I mean awesome-react, tldr-pages or homebrew-cask are probably not unimportant but that's not where I contributed most to.
- Phillips126 6y agoI am not a huge GitHub user and have only contributed some code to a single repo that was merged. I was surprised to see I had the badge in my profile.
- etaioinshrdlu 6y agoIt looks like the code is actually stored in plain text, and that this is basically microfilm?
- rob-olmos 6y agoI don't think so. Project Silica talks about storing the data in droplet-looking voxels rather than etching language symbols. Cool video of the process: https://www.youtube.com/watch?v=6CzHsibqpIs https://www.youtube.com/watch?v=6CzHsibqpIs
- etaioinshrdlu 6y agoBut, from the article, it doesn't look like they used Project Silica here, they used piqlFilm.
- rejschaap 6y agoYeah, Project Silica is another project within the GitHub Archive Program. You can see the microfilm in this video https://www.youtube.com/watch?v=fzI9FNjXQ0o&feature=youtu.be&t=72 https://www.youtube.com/watch?v=fzI9FNjXQ0o&feature=youtu.be...
- rezendi 6y agoArchive Program director here. It is basically microfilm (albeit very long-lived) but the data is mostly stored in a pixellated form, not unlike QR codes, although every reel also contains human-readable instructions (and code) re how to unpack its data.
- tw4l 6y agoAs David Rosenthal (formerly of Sun, NVIDIA, and Stanford) explains, the actual Arctic Code Vault is a PR stunt, and has almost no chance of helping anyone in any kind of realistic disaster scenario: https://blog.dshr.org/2019/11/seeds-or-code.html https://blog.dshr.org/2019/11/seeds-or-code.html That said, the rest of the project, which focuses on preserving several independent copies of repositories hosted on GitHub with a handful of partner organizations, is quite useful. From the same post: "They are using a range of technologies, making feeds available over the Internet, and partnering with the Internet Archive, the Software Heritage Foundation and the Bodleian Library. These are mostly things which will get used in the foreseeable future, and should be applauded for that reason."
- Wowfunhappy 6y ago>> They drag the 200 platters out into the 24hr sunshine, plug the solar panel into the Raspberry Pi, point its camera through a magnifying glass at the first frame, and let the QR app they happen to have on the Pi's micro--SD card do its thing. A couple of seconds later they have the first 2,900 bytes on the USB drive. It takes another couple of seconds to move to the next frame by hand. So they sit there for 383 days scanning a frame every 4 seconds to decode the entire archive. Except there's only sunshine enough for the Pi half the year, so it takes rather more than two years. Then they need to start the Pi building all that code... >> Of course, this is ridiculous. No-one will decode this archive in the foreseeable future. Yes, no one will be digging code out of Github right after the apocalypse. But what about 200 years after the apocalypse? Or maybe just 1,000 years from now, no apocalypse needed? I could see the archive being of immense historical value.
- kevin_thibedeau 6y agoThe Pi isn't going to work after 200 years. Its flash will be wiped. Never mind aging on all the other parts.
- jedieaston 6y agoPresumptively, the Tech Tree will have some way of bootstrapping a system capable of decoding the tapes. They say in the introduction that it’s nearly useless to access the tapes without a computer and that they expect whoever is reading this is to have a computer that is centuries more advanced than we have now. Maybe they just zip tied a ThinkPad to the tape reader and pray that it can eat whatever happens to it in the vault.
- nomoreservices 6y agoStrong A Fire Upon the Deep vibes thinking of future archeologists studying that.
- atonse 6y agoThis is so awesome, but the most surprising to me is that all the public source code on GitHub only totals 21 TB. I forget that they do fundamentally host text, and not video etc. I somehow thought it would be petabytes. The private repos might be more than that but those are historically paid.
- decko 6y ago20 of those are probably node_modules folders
- tuananh 6y agonode_modules wouldn't make it to git repo. at least, the top 6000 repo on github. that's for sure.
- no_wizard 6y agoOn the topic of size, I wonder how small it would be if you were able to deduplicate all repositories against each other. I sometimes suspect there is a tremendous amount of copy/paste code out there masquerading as someone else’s. Even a naive deduplication might yield some very interesting results Reminds me of a time I caught someone using someone else’s code in an interview and passing it off as their own. (Using was fine, it was the claim that it was theirs that bugged me)
- progval 6y agoI work at Software Heritage, where we archive all source code we can find, including all GitHub repositories, and deduplicate them internally. The size of all file contents (including older versions of files) is a few hundreds TBs, and everything else (directory structures, revision history, etc.) is under 10TB. So for GitHub alone it would be a little under that
- 1337shadow 6y agoThey've just archived the HEAD of the 6000 most popular repos > We’ve archived 6,000 of the world’s most popular repositories as a proof of concept for future archives. > The snapshot will consist of the HEAD of the default branch of each repository, minus any binaries larger than 100KB in size.
- girst 6y agoIn a 1000 years people will surely benefit from the millions of copy-pasted dotfiles :^)
- rwky 6y agoThis means that after the apocalypse people will be able to reclaim the Linux source code but not Windows. I find it poetic that open source may one day be the norm.
- xaedes 6y agoSince Microsoft aquired Github, I think they may also put some MS closed-source code in the vault. Seperate.
- userbinator 6y agoThere are copies of leaked Windows source code floating around... I've even seen it on GitHub but they probably get DMCA'd pretty quickly.
- jedieaston 6y agoI’m thinking that someone at Microsoft may have snuck the code for Windows into the archive after it was pulled from Github. Between Windows and OS X, a ton (most?) of the end user software would be unusable to a future generation in its original form since they didn’t have the desktop OS it was used on. Ironically, 500 years from now, they may think that the year of the Linux desktop was 2008 :-D
- zaptrem 6y agohttps://github.com/reactos/reactos https://github.com/reactos/reactos would probably make this less of an issue as well.
- symplee 6y ago1000 years from now, I can only imagine the hidden Y3K bugs...
- ca_parody 6y agoHonestly, for however much this project either (a) is a genuine archeological move for the preservation of information or (b) to get good press, all I genuinely thought when this happened is "aw shucks - wish i fixed those bugs before they zapped it onto film and flew it to santa clause".
- jcahill 6y agoI am a web archivist with an archival project on Svalbard that predates this GitHub initiative. Additionally, large-scale github-specific projects like https://gharchive.org https://gharchive.org (formerly GitHub Archive) have existed for some time. In my experience, code is more likely than not to be preserved in a stale revision, if at all. The most common forms of preservation are (a) simple tarballing and (b) git bundles.
- benatkin 6y agoIce. Not to be confused with ICE.
- rvz 6y agoGitHub is working with both? Very chilling.
- Google234 6y agoThis is a waste of money.
- fnord77 6y agohttps://en.wikipedia.org/wiki/5D_optical_data_storage https://en.wikipedia.org/wiki/5D_optical_data_storage
- un_montagnard 6y ago> The next morning, it traveled to the decommissioned coal mine set in the mountain, and then to a chamber deep inside hundreds of meters of permafrost, where the code now resides fulfilling their mission of preserving the world’s open source code for over 1,000 years. What is the probability that we still have the required tech to read that code in 1,000 years?
- davedx 6y agoDepends on whether the Great Filter is before or behind us.
- TheSpiciestDev 6y agoHa, that README grammar fix years ago finally pays off!
- grogenaut 6y agoWhat I really want from github is to allow people who own open source projects who don't want to own them anymore to just hand them off for escrow so that at a later date a reputable group like apache can maintain them if needed.
- sudhirj 6y agoCouldn’t Apache just make a fork and announce it? Or is this just about the convenience and marketing?
- grogenaut 6y agoIf it's done this way then all of the web links stay live, and a new owner doesn't need to be found immediately. Think of it as a special permission holding pool. There are many cases of "done" libraries that need changes later. This would help with them. However when they're not done this way you can spend a few weeks / months trying to get ahold of the author and for them to decide "oh yeah I don't really care about x anymore"
- onion2k 6y agoIf someone wants to give up their project and there isn't anyone in the community who wants to take over, the project is already dead. Open source doesn't work without people around to push it.
- 1337shadow 6y agoWhere can we find the list of the 6000 repos ? On my profile it just shows 3 "and more", would like to get the full list. TYIA ;)
- tazard 6y agoI'm really curious about this too. I haven't been able to find this information anywhere.
- deleted 6y ago[deleted]
- axegon_ 6y agoSame. Or how they were picked. I kept scratching my head all evening cause I haven't made any updates or contributions to mine in quite a while.
- zenhack 6y agoMy best guess is it's some function of the popularity. The three that my profile shows are - capnproto/capnproto - sandstorm-io/sandstorm - erlang/otp (I don't remember the order). I actively contribute heavily to sandstorm. I've sent patches here and there to capnproto, and it's vaguely a sister project to sandstorm. Those are probably some of the most popular projects I have multiple contributions to, though there are others. otp feels a bit odd though, if there's and "and more" -- I sent them a one line patch to fix a build error when building against musl. I haven't really been involved since, nor was I before. But it's a high profile project.
- tuananh 6y agowhere did you get that 6000 repos number?
- deleted 6y ago[deleted]
- 6y ago
- juanbyrge 6y agoThis is pointless - a complete waste of time, effort, and energy. Isn’t there something more beneficial they could have done instead? Why pollute the Arctic with plastic and film canisters?
- malechimp 6y agoIf things come to that I doubt the practicality of it all. But it makes easy headlines. It also makes open source an immortality project for a lot of people.
- deleted 6y ago[deleted]
- Zamicol 6y agoThey __are not__ using QR code for storage as has been misreported by a few media outlets. See https://earth.esa.int/documents/1656065/3222865/170922-Piql-ESA_Slides-Final https://earth.esa.int/documents/1656065/3222865/170922-Piql-... for piql's storage method.
- deleted 6y ago[deleted]