5 ms·
One of the important avenues to scrape AJAX heavy and phantomjs avoiding websites is using the google chrome extension support. They can mirror the dom and send
by bootcat 9y ago
One of the important avenues to scrape AJAX heavy and phantomjs avoiding websites is using the google chrome extension support. They can mirror the dom and send it to an external server for processing where we can use python lxml to xpath to appropriate nodes. This worked for me to scrape Google, before we hit the capatcha. If anyone is interested, i can share code i wrote to scrape websites !
If you can scrape findthecompany database ? I have done it successfully !!
- visarga 9y ago> This worked for me to scrape Google, before we hit the capatcha. If Google wanted to give back something to the community, it would offer cheap automated searches (current prices are absurd). Another thing - more depth after the first 1000 results. Sometimes you want to know the next result. We shouldn't need to do all these stupid things to batch query a search engine, it should be open. That makes it all the more important to invent an open-source, federated search engine, so we can query to our heart's content (and have privacy).
- bootcat 9y agoI absolutely agree, and I am thinking strategies to even automate the capatcha, using crowdsourcing or better, using AI/ML ( which is not trivial ). duckduckgo is good but not there yet. Would you be interested to work on a search engine ? Some projects are bitfunnel and so forth.
- zapperdapper 9y agoAgree 100% too. As for 'federated search engine' - it's not 'federated' per se but check out Gigablast search engine. Open source (source on GitHub) and a TOTALLY AWESOME piece of software written by one guy. You can do good searches at the Gigablast site[1], or set up your own search engine. Gigablast also offers an API (I may be wrong but I think DuckDuckGo uses that API for some tasks). [1] http://gigablast.com http://gigablast.com