5 ms·
I love the New York Times shots. It's a great demonstration of how off-putting their interstitial ads are, and how many other sites don't need them.
by there 14y ago
I love the New York Times shots. It's a great demonstration of how off-putting their interstitial ads are, and how many other sites don't need them.
- palewire 14y agoHa! If you know a way around them I'd appreciate the tip.
- pavel_lishin 14y agoHow are you fetching them now?
- palewire 14y agoJust a stoopid straight call with Selenium's Python bindings.
- knowtheory 14y agoI'd think that if you registered a user and submit the session w/ the user cookie you'd probably never hit the interstitial
- donohoe 14y agoHmm. Not sure if it would make a difference but maybe: (1) Try this URL instead: http://www.nytimes.com/pages/ http://www.nytimes.com/pages/ (2) Hit the URL by date: http://www.nytimes.com/indexes/yyyy/mm/dd/ http://www.nytimes.com/indexes/yyyy/mm/dd/ Example: http://www.nytimes.com/indexes/2011/12/03/ http://www.nytimes.com/indexes/2011/12/03/ You can also use this URL structure to get the Homepage back several years (2001) as it was around midnight of that date. http://www.nytimes.com/indexes/2001/01/01/ http://www.nytimes.com/indexes/2001/01/01/ http://www.nytimes.com/indexes/2001/09/11/ http://www.nytimes.com/indexes/2001/09/11/ (Notable) Not sure if this (or any other section) is of interest too: http://www.nytimes.com/indexes/2010/12/03/todayspaper/ http://www.nytimes.com/indexes/2010/12/03/todayspaper/ (3) When you scrape the page find out the link it provides to the Homepage and then try that. I had some success doing that. What I really want is THIS: "Reward - NYTimes Login Script" http://donohoe.tumblr.com/post/10723388191/reward-nytimes-login http://donohoe.tumblr.com/post/10723388191/reward-nytimes-lo... which would get around that problem.
- schwanksta 14y agoWoah, the /indexes/ url just blew my mind; so it just gives you the last update on that date?
- donohoe 14y agoYes. NYT homepage editor has flexibility on when they "roll" the date depending on news events (and content can be ranked ahead of time too). But yeah, basically midnight. Internally I know editorial preserves them on an hourly basis. You might ask current NYT-ers if you could get access to that? I know I did a scrape every 10 (or 2?) minutes of the Homepage HTML for a couple of years (2007 to 2011?). If I can find that data and its still meaningful I'll get it to you. It was quite a few GB.
- palewire 14y agoThanks for this great information. I'm stuck at the jury duty cattle call this morning but will try to put this into action later.
- jashkenas 14y agoTake care with putting this into action -- the external homepage archives constantly suck in the latest version of any precoded module on the page ... so for example, this should show the Iowa Caucuses, but shows nearly two months later instead: http://www.nytimes.com/indexes/2012/01/03/ http://www.nytimes.com/indexes/2012/01/03/ ... a misstep we later fixed. For what it's worth, there's also an internal version of the homepage archive that doesn't suffer from this problem, and is snapshotted hourly.
- danso 14y agoSo even all the external stylesheets/js are (relatively) preserved? That's pretty slick...was this something that was retroactively applied, or something that's existed since early iterations of the CMS?
- adventureful 14y agoIt seems like the NYT doesn't show the interstitial ads on immediate repeat visits (once per day perhaps?). Hit the nyt.com site, then hit it again shortly thereafter and snap that visit. Can your approach accept cookies?