7 ms·
Using Node.js and JQuery to Crawl Public Tweets
- blyxa 14y agowhy not use the twitter api?
- bdreadz 14y agofrom the github page: Birdeater does not use Twitter's API. It was built as a demonstration of an approach I like to use for parsing structured information from unstructured HTML.
- TazeTSchnitzel 14y agoA better (and practical) example is scraping an internet forum (I've done it, partially)
- sshillo 14y agoapi has rate limits
- dbau 14y agoSo does this it seems: Be mindful when running it, as Twitter limits the number of requests that a single client can make per hour.
- wamatt 14y agoThe API probably has a key that can be blocked. Not that I'm advocating this, but a potential advantage of scraping, is it can be combined with tor or proxies, etc to get around limits.
- dmazin 14y agoIt's rate-limited to 150 requests per hour and you can only pull the 3,200 most recent tweets for a user.
- hafabnew 14y agoFrom the docs: ''' * Node.js [...] * jQuery [...] [...] This approach has become my hammer when web scraping tasks come up. ''' If all you have is a hammer, you may find yourself noticing that objects become more nail-like :).
- latchkey 14y agoIf you really want to scrape pages, you should use something like https://github.com/chriso/node.io/ https://github.com/chriso/node.io/ which batches things in jobs, helps with error handling, io, etc...
- danso 14y agoDoes Node have anything like Mechanize? Handling cookie state and such is something that is much more useful than the selector functionality of jQuery...which is great, but not any better than what Nokogiri offers.
- thegoleffect 14y agohttps://github.com/sgentle/phantomjs-node https://github.com/sgentle/phantomjs-node is pretty good for most tasks.
- laughinghan 14y agohttp://zombie.labnotes.org/ http://zombie.labnotes.org/ is a library I've used with great success. The documentation in particular is cute. I found PhantomJS unnecessarily convoluted for trivial tasks and was unable to figure how to do the nontrivial thing I was actually trying to do. The documentation in particular was unusable.
- wavephorm 14y agoUsing JQuery server-side to process Twitter posts which are already in JSON format is just so dumb I can believe I read an entirely usless blog post about it.
- wskinner 14y agoI have also found node+jQuery an effective web crawling combination. In particular the cheerio library https://github.com/MatthewMueller/cheerio https://github.com/MatthewMueller/cheerio greatly simplifies data extraction. And as others have mentioned, the asynchronous nature of node is perfectly suited to crawling (as long as you take care not to accidentally DDOS the target site).