3 ms·
I don't quite understand why you would use a full-blown browser like phantomjs for crawling (I've seen a lot of projects recently taking this approach, so this
by thecodemonkey 11y ago
I don't quite understand why you would use a full-blown browser like phantomjs for crawling (I've seen a lot of projects recently taking this approach, so this critique is not directly towards Apifier).
Yes, I get that in some specific circumstances it would be nice to be able to execute the JavaScript on the page but think about the trade-off here.
In the vast majority of cases a simple HTTP GET request with a DOM parser is all you need -- actually not a single one of the examples on the Apifier homepage has any need for phantomjs.
Wouldn't it be much much cheaper, simpler and faster to ditch phantomjs? Or is there something I'm missing here?
- jancurn 11y agoYou're right that most of the time you don't need to use JavaScript. But look at Google Groups for example - there is an infinite scroll to get all the topics, posts are also loaded dynamically, so you have to wait some time to get them. In the SFO flights example you have to deal with pagination also using JavaScript. We wanted to build a powerful tool which can crawl and scrape almost any website out there. It's slower, but you can use bench of our nodes to do it in parallel.
- deleted 11y ago[deleted]
- thecodemonkey 11y agoI agree that the Google Groups example is much simpler when using PhantomJS, but I would argue that it would be an outlier. The SFO flights example is actually heavily over engineered, from quickly glancing over the XHR tab in Chrome Network tools it was pretty obvious that all of the data is actually located in this very nice JSON blob http://www.flysfo.com/flightprocessing/fullFlightData.txt http://www.flysfo.com/flightprocessing/fullFlightData.txt (I assume that the SF Flight Info was just meant as an example for the platform and as such the fact that it's already a JSON blob was just ignored for the sake of the example)
- Eridrus 11y agoI've written some projects which use phantomjs; the primary motivation for me has been the desire to look at the web in general, rather than specific sites I'm scraping data off, and having the ability to see what their javascript does.
- est 11y agoIt's OK when you only have to crawl one or two websites, sure, manually analyze the js and write minimal DOM parsing routes would do. But how about hundreds or thousands of websites to crawl? Or do you prefer just use phantomjs write static extraction rules.
- pkulak 11y agoIf you limit to just that, then there's no benefit over 10 lines of Ruby and Nokogiri.