4 ms·
Do you have a rough set of guidelines for how fast we should request from HN? For a side project, I was thinking of writing something that scraped the HN front
by someone13 14y ago
Do you have a rough set of guidelines for how fast we should request from HN? For a side project, I was thinking of writing something that scraped the HN frontpage and all the associated comment threads every 10 minutes or so, and I'd rather not cause performance issues or get banned. I'd be happy to rate-limit requests to whatever is convenient.
- citricsquid 14y agoThis might be helpful: http://api.ihackernews.com/ http://api.ihackernews.com/ edit: oh, official API is above. Disregard this one :-)
- unreal37 14y agoMay be better to use the official API. http://www.hnsearch.com/api http://www.hnsearch.com/api
- tallanvor 14y agoIf it were an official API, wouldn't it be associated with HN or Y Combinator rather than an external website?
- wglb 14y agoIt is by a YC company, and recommended by pg.
- laumars 14y agoThat's not an official API: http://www.hnsearch.com/about http://www.hnsearch.com/about Quote: "HNSearch was built by the team at ThriftDB to give back to the community and to test the capabilities of the ThriftDB flexible datastore with search built-in." Interesting API all the same though.
- zargon 14y agoRegardless whether it is official or not, it is pg's preferred api: http://news.ycombinator.com/item?id=4694308 http://news.ycombinator.com/item?id=4694308
- mvanveen 14y agoThe robots.txt file for HN suggests a Crawl-Delay value of 30 seconds.