7 ms·
Meet PastPages.org, the news homepage archive. (And help keep it alive)
I created this site because I think it ought to exist. The shifting homepages of major media sites should be saved so they can be studied. Done right, I believe PastPages could serve as a resource for scholars seeking to study coverage of news events, like the upcoming U.S. presidential election.
Regularly collecting the data costs money. So I've organized a Kickstarter in hopes of raising funds to keep it up. http://www.kickstarter.com/projects/651552740/keep-pastpages-alive
- palewire 14y agoI created this site because I think it ought to exist. The shifting homepages of major media sites should be saved so they can be studied. Done right, I believe PastPages could serve as a resource for scholars seeking to study coverage of news events, like the upcoming U.S. presidential election. Collecting this data cost money. So I've set up a Kickstarter drive to raise funds. If you'd like to help keep PastPages alive, please considering giving. http://www.kickstarter.com/projects/651552740/keep-pastpages-alive http://www.kickstarter.com/projects/651552740/keep-pastpages...
- atlbeer 14y agoWhat stack are you using for the web page capture? It's a perfect crisp capture. I've tried before and never got such good programatic results.
- palewire 14y agoI'm using Selenium's Firefox driver from inside a Django app. There is Python binding that's slick once you figure out a couple timeout related workarounds that are necessary. Their forums helped me over that hurdle.
- showerst 14y agoDo you have permission to be taking these snapshots? The Newseum does something similar for print content and has agreements with all of the organizations so that they don't just Cease & Desist them out of existence. http://www.newseum.org/todaysfrontpages/ http://www.newseum.org/todaysfrontpages/
- kgen 14y agoDoesn't Fair Use cover "commentary, criticism, news reporting, research, teaching, library archiving and scholarship."? This site seems non-profit and educational in nature, so I would imagine it would pass the balancing test if push came to shove.
- there 14y agoI love the New York Times shots. It's a great demonstration of how off-putting their interstitial ads are, and how many other sites don't need them.
- palewire 14y agoHa! If you know a way around them I'd appreciate the tip.
- pavel_lishin 14y agoHow are you fetching them now?
- palewire 14y agoJust a stoopid straight call with Selenium's Python bindings.
- knowtheory 14y agoI'd think that if you registered a user and submit the session w/ the user cookie you'd probably never hit the interstitial
- donohoe 14y agoHmm. Not sure if it would make a difference but maybe: (1) Try this URL instead: http://www.nytimes.com/pages/ http://www.nytimes.com/pages/ (2) Hit the URL by date: http://www.nytimes.com/indexes/yyyy/mm/dd/ http://www.nytimes.com/indexes/yyyy/mm/dd/ Example: http://www.nytimes.com/indexes/2011/12/03/ http://www.nytimes.com/indexes/2011/12/03/ You can also use this URL structure to get the Homepage back several years (2001) as it was around midnight of that date. http://www.nytimes.com/indexes/2001/01/01/ http://www.nytimes.com/indexes/2001/01/01/ http://www.nytimes.com/indexes/2001/09/11/ http://www.nytimes.com/indexes/2001/09/11/ (Notable) Not sure if this (or any other section) is of interest too: http://www.nytimes.com/indexes/2010/12/03/todayspaper/ http://www.nytimes.com/indexes/2010/12/03/todayspaper/ (3) When you scrape the page find out the link it provides to the Homepage and then try that. I had some success doing that. What I really want is THIS: "Reward - NYTimes Login Script" http://donohoe.tumblr.com/post/10723388191/reward-nytimes-login http://donohoe.tumblr.com/post/10723388191/reward-nytimes-lo... which would get around that problem.
- xabi 14y agoSame service here, but for newspapers (with more than 1000 newspapers around the globe): http://en.kiosko.net/us/ http://en.kiosko.net/us/
- sp332 14y agoI think hosting a static image is cool, but doesn't the Internet Archive already have a full-HTML archive of these pages? e.g. http://web.archive.org/web/20110729013424/http://www.nytimes.com/ http://web.archive.org/web/20110729013424/http://www.nytimes...
- palewire 14y agoIt does an I love that site but unless I'm mistaken I don't think they grab often enough to track fast moving news events.