7 ms·
Any good api to scrape HN other than this?
how to scrape HN other then https://github.com/karan/HackerNewsAPI . any good premade library in python ?
- culo 13y agotry these - https://www.mashape.com/scrape/scrape-it#!documentation https://www.mashape.com/scrape/scrape-it#!documentation - https://www.mashape.com/karangoel/hnify#!documentation https://www.mashape.com/karangoel/hnify#!documentation
- notastartup 13y agohaha good to see someone link it! I am the author of Scrape.it currently on mashape. I also wrote http://scrape.ly http://scrape.ly for crawling web pages and extracting data.
- napoleond 13y agoJust use https://www.hnsearch.com https://www.hnsearch.com, along with https://www.hnsearch.com/rss https://www.hnsearch.com/rss and https://www.hnsearch.com/bigrss https://www.hnsearch.com/bigrss if you want to mimic the front page. There is rarely a need to scrape HN directly, but if you do make sure your bot is polite (especially with respect to rate limits).
- kaushikfrnd 13y agoI am trying to fetch all posts,comments plus all user data . I will ty hnsearch .
- jenjenhar 13y agoOut of curiosity, Why does HN not release an official API?
- code_duck 13y agoMy impression is that pg wants to encourage the hacker spirit by providing a bare bones service which could easily have a 'hacked' api built upon it.
- taliesinb 13y agoMy impression is that HN's link and comment data is too valuable for pg to give away. Certainly, if I have had access to it I know I could do some pretty useful sociology on HN's audience (= the pool of startup hire material).
- code_duck 13y agoI don't believe that HN restricts or discourages the scraping of HN content in any way... Other than the restrictions here: https://news.ycombinator.com/robots.txt https://news.ycombinator.com/robots.txt If you have a fabulous idea for how to use the data contained on this site, I'm sure everyone will be impressed and interested to see it.
- kaushikfrnd 13y agoi had the same question in my mind . Even reddit have there official api .
- toyg 13y agoI bet it's just a cost/benefit analysis. An API is a way to get more eyeballs by motivating 3rd party developers to integrate and publicise your service. HN does not need that: it has enough traffic as it is, and given the target audience, you would see an instant proliferation of half-assed apps hammering its endpoints. So it would be an additional cost for no real benefit. The current situation (PG and friends optimise a basic but very accessible website, and a handful of third parties build APIs on top) is much more manageable.
- carbocation 13y agoThe robots.txt from news.ycombinator.com reads as follows: User-Agent: * Disallow: /x? Disallow: /vote? Disallow: /reply? Disallow: /submitted? Disallow: /submitlink? Disallow: /threads? Crawl-delay: 30 So nominally you should feel free to set up a scraper that crawls one non-disallowed resource every 30 seconds.
- t0 13y agoBut /x? is for the next page.
- pedrocr 13y agoSo apparently you can get two pages of ranking, using / and /news2.
- randomdata 13y agoDepends on the intent. If it is user-initiated (like say a mobile formatted version of the site), it wouldn't have to be obey the robots.txt, since it is not a crawler, just another web browser.
- kaushikfrnd 13y agowell i am trying to get user submissions also so may be i have to violate the robots.txt
- nashequilibrium 13y agorp = robotparser.RobotFileParser() rp.set_url("https://news.ycombinator.com/news/robots.txt" https://news.ycombinator.com/news/robots.txt") rp.read() # Reads the robots.txt rp.can_fetch("*", 'https://news.ycombinator.com/news' https://news.ycombinator.com/news') >>>> True
- kaushikfrnd 13y agocool
- mvanveen 13y agoI have a ScraPy-based crawler project available at http://github.com/mvanveen/hncrawl http://github.com/mvanveen/hncrawl
- jcla1 13y agoNot a full featured api, but a way to scrape all of HN: http://jcla1.com/blog/2013/05/13/crawling-hackernews/ http://jcla1.com/blog/2013/05/13/crawling-hackernews/ Disclaimer: It's my own blog edit: Uses HNSearch, so it doesn't violate the robots.txt and can be crawled faster
- zerd 13y agoDid you manage to download the whole database that way? Edit: Also, why didn't you use the "start" (offset) parameter?
- jcla1 13y agoNo, not tried to download it yet. Regarding your question, if you try to use a start > 999 you get this error: "Validation error: max limit is 100, max start+limit is 1000", which is why I avoided that parameter.
- goldenkey 13y agoYahoo pipes would work really well if you're willing to write a few HTML regexes or dom element selectors. http://pipes.yahoo.com/pipes/ http://pipes.yahoo.com/pipes/
- deft 13y agoI wrote an alright one in Python for use in my HN app for BlackBerry 10. Not sure how good it is, but check it out here: https://github.com/krruzic/Reader-YC/tree/master/app https://github.com/krruzic/Reader-YC/tree/master/app I'm not sure what you're trying to do though. I used beautifulsoup because I couldn't get lxml working on BB10, but if it was switched to using lxml it would be much faster.
- dmpayton 13y agoI wrote a Python wrapper for the iHackerNews API, if that helps. https://github.com/dmpayton/python-ihackernews https://github.com/dmpayton/python-ihackernews
- kaushikfrnd 13y agoi saw your github repo . Wonderful work but saw your api was not working getting some errors when i tried the link http://api.ihackernews.com/by/kaushikfrnd http://api.ihackernews.com/by/kaushikfrnd. Can you confirm it will work if i run it on my own server .
- dmpayton 13y agoAh, looks like there's an issue with the iHackerNews API itself, which I don't have a hand in. You'll want to hit up @ronnieroller on Twitter. Sorry I can't be of more help. :/
- droid_w 13y agoThere's a twitter feed based on HN - https://twitter.com/newsycombinator https://twitter.com/newsycombinator You can use the twitter API and read from there
- mikektung 13y agoDepending on what you're trying to do with the data, you may find http://diffbot.com/products/automatic/ http://diffbot.com/products/automatic/ helpful for getting the clean article text and categorization in JSON format. It can be used as a complement/augmentation to the great suggestions here for getting the links. Disclosure: Founder of Diffbot here.
- notastartup 13y agoI wrote http://scrape.it http://scrape.it and http://scrape.ly http://scrape.ly to do this.
- cheeaun 13y agoI built https://github.com/cheeaun/node-hnapi https://github.com/cheeaun/node-hnapi
- amirouche 13y agoThere is hundred of data sets out there why it must always be HN?
- zerd 13y agoBecause quality datasets are hard to get. E.g. on reddit you would just get cats and memes.
- obayesshelton 13y agoYou don't even need an api, all you need is an rss reader and read - https://news.ycombinator.com/rss https://news.ycombinator.com/rss
- fakename 13y agoother than this
- kaushikfrnd 13y agocan anyone say me how to get https://news.ycombinator.com/news https://news.ycombinator.com/news through hnsearch api . I want the api link -> [http://api.thriftdb.com/api.hnsearch.com/ http://api.thriftdb.com/api.hnsearch.com/] !!
- rotub 13y agohttps://www.hnsearch.com/api https://www.hnsearch.com/api
- shamsulbuddy 13y agohttp://hnapp.com/ http://hnapp.com/ -- This is the best HN Scraped site.. returns data in JSON / RSS format.