4 ms·
HN2JSON: A ruby gem for HackerNews
- qmacro 14y agoExcellent! I know I'm biased but I also know you've put a lot of effort into this. Well done Joseph.
- markburns 14y agoitem = HN2JSON.find 4623690 NoMethodError: undefined method `url=' for #<HN2JSON::Entity:0x007fb84cd63a88> from /Users/markburns/.rvm/gems/ruby-1.9.3-p194/gems/hn2json-0.0.4/lib/hn2json/parser.rb:92:in `block in get_attrs_post' from /Users/markburns/.rvm/gems/ruby-1.9.3-p194/gems/hn2json-0.0.4/lib/hn2json/entity.rb:92:in `add_attrs' from /Users/markburns/.rvm/gems/ruby-1.9.3-p194/gems/hn2json-0.0.4/lib/hn2json/parser.rb:91:in `get_attrs_post' from /Users/markburns/.rvm/gems/ruby-1.9.3-p194/gems/hn2json-0.0.4/lib/hn2json/entity.rb:71:in `get_attrs' from /Users/markburns/.rvm/gems/ruby-1.9.3-p194/gems/hn2json-0.0.4/lib/hn2json/entity.rb:56:in `initialize' from /Users/markburns/.rvm/gems/ruby-1.9.3-p194/gems/hn2json-0.0.4/lib/hn2json.rb:35:in `new' from /Users/markburns/.rvm/gems/ruby-1.9.3-p194/gems/hn2json-0.0.4/lib/hn2json.rb:35:in `find'
- dfc 14y agoBe careful not to hammer the site. Your IP could be added to the blocklist if you are too aggressive: "Yes, we block IPs that seem to be crawlers ignoring robots.txt. We've always blocked abusive IPs, but I tightened up the blocking a few weeks ago. A lot of people were crawling HN, most of them unnecessarily because they were doing things they could have done more efficiently through HNSearch's API[1]." --pg[2] [1] http://www.hnsearch.com/api http://www.hnsearch.com/api [2] http://news.ycombinator.com/item?id=3196298 http://news.ycombinator.com/item?id=3196298
- rdudekul 14y agoGoing through the code on github to see how a HN page is parsed, was informative. I may use this to create one using Node.js. My interest is in building an intelligent agent that filters content based on my interests (example: coding, customer acquisition, hiring etc.) and notifies me on a daily or weekly basis.
- jcla1 14y agoI have written a program that is similar to what you just explained, also on GitHub https://github.com/jcla1/hackernews https://github.com/jcla1/hackernews
- why-el 14y agoNice work. Does Cronic have to be a runtime dependency?
- jcla1 14y agoNot really, but at the time I didn't want to have to write my own date parser. (HN doesn't show the date, just things like "x days ago")
- mmackh 14y agoI've written a script that extracts HN, which anyone is welcome to use. I use it for the Hacker News iPhone app: http://api.thequeue.org/hn/frontpage.xml http://api.thequeue.org/hn/frontpage.xml http://api.thequeue.org/hn/new.xml http://api.thequeue.org/hn/new.xml http://api.thequeue.org/hn/best.xml http://api.thequeue.org/hn/best.xml
- mmackh 14y agoI've written something similar for my Hacker News iPhone app: http://api.thequeue.org/hn/frontpage.xml http://api.thequeue.org/hn/frontpage.xml http://api.thequeue.org/hn/new.xml http://api.thequeue.org/hn/new.xml http://api.thequeue.org/hn/best.xml http://api.thequeue.org/hn/best.xml
- mvanveen 14y agoI wrote a small, ScraPy based HN crawler available at http://github.com/mvanveen/hncrawl http://github.com/mvanveen/hncrawl in case anyone is interested.
- selvan 14y agoCheckout apify - http://apify.heroku.com/resources http://apify.heroku.com/resources & scrapify - https://github.com/sathish316/scrapify https://github.com/sathish316/scrapify Library to scrap HTML content as JSON data.