3 ms·
With hnsearch.com being shutdown, I believe he is referring to: http://hn.algolia.com/api http://hn.algolia.com/api
by jared314 13y ago
With hnsearch.com being shutdown, I believe he is referring to: http://hn.algolia.com/api http://hn.algolia.com/api
- sillysaurus3 13y agoUsing that API, you could download 48,000 hacker news stories in 2 days, so if there have been less than 48,000 submissions, then what minimaxir said is true. But first you'd need to generate a list of all 48,000 story ids, and there seems to be no way to actually do that.
- minimaxir 13y agoNo need to generate all IDs beforehand. The search_by_date endpoint is fine. You have to paginate using the created_at_i parameter, not the page parameter. Also, you can set hitsPerPage = 1000. ;)
- sillysaurus3 13y agoWhen trying to access page 2 via that endpoint: http://hn.algolia.com/api/v1/search_by_date?tags=story&hitsPerPage=1000&page=2 http://hn.algolia.com/api/v1/search_by_date?tags=story&hitsP... "you can only fetch the 1000 hits for this query, contact us to increase the limit" It was a nice try, but it did seem too good to be true.
- minimaxir 13y agoYou can paginate using the created_at_i parameter (edited OP) Just pass created_at_i<X, where X is the time stamp of the earliest submission. I was able to download 500k stories (i.e. about half of HN's 1.26M stories) before I ran into memory issues; I've fixed them and am downloading the rest.
- sillysaurus3 13y agoThat's incredible. Will you upload the raw database somewhere, please? If you make a torrent, I'll help seed it. Would you email me at sillysaurus3@gmail.com whenever it's ready?
- deleted 13y ago[deleted]