4 ms·
Ask HN: Does anyone have a backup of Aaron Swartz' site: theinfo.org?
Hi HN folks,
I was browsing Aaron's blog [http://www.aaronsw.com/weblog/rawnerve http://www.aaronsw.com/weblog/rawnerve] when I noticed the footer links out to http://theinfo.org/ http://theinfo.org/ (last updated March 2008). However, none of the links are still alive, and I couldn't find any pages that were backed up on Archive.org.
I figured this would be the best place to ask - does anyone have an archive of these pages?
- deadalus 5y agohttps://web.archive.org/web/20090714025054/http://theinfo.org/get/data https://web.archive.org/web/20090714025054/http://theinfo.or...
- sillysaurusx 5y agoHoly hell. It's incredibly inspiring what aaronsw was able to do. It's a list of links, and I kept scrolling expecting to run into text. Nope. Just a gigantic list of links, and choosing a random one seems to result in actual data: https://web.archive.org/web/20090410081645/http://www.rdfabout.com/demo/census https://web.archive.org/web/20090410081645/http://www.rdfabo... Well, almost actual data. The 4.7MB tarball link is broken, but the author notes: > For the detailed Census statistics, you'll have to download the raw Census data files from the Census Bureau, my Perl script and the patch file below and run it yourself because the files are too big for me to offer as a download! And the perl script link still works. From experience, I know how valuable that can be. The only reason I was able to help build The Pile (books3) is because of aaronsw's html2text script (https://github.com/aaronsw/html2text https://github.com/aaronsw/html2text). Out of at least four conversion options I tried, that script was the only one that was flexible enough to be modified to spit out human-readable text at scale. Thanks for the nostalgia. 2008 doesn't feel like yesterday anymore, but it will always feel incredibly special. That era was just different, and it's a shame that one only realizes it in hindsight. Otherwise I would've slowed down just to look around more and appreciate getting a glimpse. P.S. although most of the data download links are dead, there is a trick to recover some of them: try to access the links on the live site. archive.org doesn't always snag tarballs, but occasionally (very occasionally) the tarballs survive till today. You should also check the parent site itself, not just the direct url to the tarball. Sometimes you'll get lucky and the site merely went through a reshuffle rather than being taken offline.
- mywacaday 5y agoI have been asked for a donation on way back and Wikipedia this morning, small donation to both. Well worth the few dollars.
- marto1 5y agowait, he actually used CKAN and now the US government is also using it !? How horrifically ironic.
- mkl 5y agoWorking backwards: Here's a snapshot from October 2021: https://web.archive.org/web/20211020224005/http://theinfo.org/ https://web.archive.org/web/20211020224005/http://theinfo.or... The updated links in it (to http://theinfo.anandology.com/ http://theinfo.anandology.com/) seem dead though. In 2014 there are error messages, and sometimes just court records: https://web.archive.org/web/20140215110438/http://theinfo.org:80/ https://web.archive.org/web/20140215110438/http://theinfo.or... The latest I can find of what looks like the original site is 2012: https://web.archive.org/web/20121003205028/http://theinfo.org/ https://web.archive.org/web/20121003205028/http://theinfo.or... The next snapshot it's replaced with the court records.
- skoocda 5y agoI guess that was my problem- I was looking up the links via anandology.com assuming that was where they'd be archived. Thanks!
- PinguTS 5y agoWorks for me: https://web.archive.org/web/20100731171624/http://theinfo.org/ https://web.archive.org/web/20100731171624/http://theinfo.or...
- qaalid 5y agoblok*37jack hack 233 real hack
- zandorg 5y agoLet's not forget jottit.com. Aaron's website to make notes. It was useful back in the day.
- comeonseriously 5y agoAnother place to ask these types of questions is on /r/datahoarder.