5 ms·
Does it support sites which require a JS enabled browser?
by r1k 10y ago
Does it support sites which require a JS enabled browser?
- aexaey 10y agoIt doesn't. To scrape (or fake-API) js-only websites you have to either: - drive a browser (firefox/chrome) via already mentioned here selenium/webdriver (potentially hiding the actual browser window into a virtual X by wrapping the whole thing with xvfb-run), - or use one of the webkit-based toolkits: phantomjs [1] or headless horseman [2]. There is also an interesting project that combines the two, i.e. it drives a Firefox (or, more precisely, slightly outdated version of Gecko) to emulate a phantomjs-compatible API. [3] phantomjs/slimerjs are pretty popular and even have tools that run on top of them, such as casperjs [4], that geared more to automated website testing, but can be quite good at scraping or fake-APIing too. [1] http://phantomjs.org/ http://phantomjs.org/ [2] https://github.com/johntitus/node-horseman https://github.com/johntitus/node-horseman [3] https://slimerjs.org/ https://slimerjs.org/ [4] http://casperjs.org/ http://casperjs.org/
- brynedwards 10y agoI recently wrote a browser-driven scraper using Nightmare[1], which uses Electron under the hood. Another option for those who prefer python is dryscrape[2], although I haven't tried it. [1] https://github.com/segmentio/nightmare https://github.com/segmentio/nightmare [2] http://dryscrape.readthedocs.io/en/latest/ http://dryscrape.readthedocs.io/en/latest/
- nikolay 10y agoDryscrape is really cool! Thanks for sharing!
- enibundo 10y agoLast time I needed something like this I used selenium. And I use requests the rest of the time.
- hendler 10y agoSame here. I use python selenium to hit a selenium server for some speed improvements. Chrome/Firefox/Phantomjs, and can inject custom javascript over the pages. Still about 10 seconds to load, render, process a page.