7 ms·
We'd love to do that but the data isn't readily available by scraping. We could continuously fetch the parent_id post until we reach the root story, but that w
by erickerr 16y ago
We'd love to do that but the data isn't readily available by scraping. We could continuously fetch the parent_id post until we reach the root story, but that would result in a lot of extra curl requests.. V2 : )
- dvk 16y agosee: http://github.com/dkeskar/hnbot http://github.com/dkeskar/hnbot /threads is not polite to robots.txt, so discussion.rb is deprecated.
- mahmud 16y agoInstead of crawling by #id, crawl by new posts from the /newest page. For each post, split it into multiple pages, setting the parent id/title that way. Not that you have to, but a future suggestion.
- jeromec 16y agoI'm guessing the info is currently scraped from the 'threads?id=username' page, but the title to each story is already there after the word 'on'.