5 ms·
I also built a rank tracker (WhooshTraffic at the time - no longer around) and we scraped the Search Result Pages as the user would see them, spoofed the sessio
by Ixiaus 10y ago
I also built a rank tracker (WhooshTraffic at the time - no longer around) and we scraped the Search Result Pages as the user would see them, spoofed the session cookie and used human captcha solvers. We would then cache the cookies generated from the captcha solve.
It was highly effective and very scalable because you can stimulate the captcha pages easily and out-of-band generate a very large cookie pool by captcha solving. Then when you go to actually scrape google you can balance your pool of IPs and cookies, chilling them before they trigger another captcha, to handle spiking demand. Contant, sustained demand was very easy to plan for.
Anyway, we stopped using the API after the third day and realized it was not only not accurate and you couldn't turn many knobs (want to scrape results as they look for a different geographical region?)
- dchuk 10y agoThis is super interesting to me, is there anything else you can share about how you approached this? In my scraping Google experience I have found roughly the same thing where once you've passed the captcha test, you can scrape a lot more. Were you scraping with real browsers or something like Mechanize/Curl? Rate limiting at all? Proxies or real servers?
- Ixiaus 10y agoWe had a pool of IPs we were leasing on our own, proxy services get abused and are poisoned. We didn't rate limit, we would just increase the size of the cookie pool if a captcha was hit, which was rare because we would scrape n-pages till a threshold was met to prevent that session from being captcha'ed so we wouldn't have to captcha solve it. We had two pools, the primary pool and the "chilling" pool, cookies near their captcha life would cool off for a few hours before returning to the active pool which behaves just like any other resource pool, every page scraped would "borrow" a cookie out of the pool, customize the encrypted location key, and make the request with a common user agent string. Scaling it was difficult but once we had it figured out, Erlang was invaluable to us and our dependence on IPs dropped once we figured out the cookie methodology. Solving captchas is cheaper than renting IPs.
- dchuk 10y agoThank you for the reply! Sorry but one last question: can you share anymore about this statement "customize the encrypted location key"?
- Ixiaus 10y agoWhen you set the location in Google it customizes the cookie with a named field that is a capital L I believe. That field is encoded or encrypted and I could never figure it out so I just constructed a rainbow table by using phantomjs to set a location and scrape the cookie out, pairing the known location value with the encrypted value so that we could customize "the location of the search".
- Ixiaus 10y agoOh and I just used a generic HTTP request client in Erlang and xpath / HTML parser to extract what we needed.
- oneeyedpigeon 10y agoGod, page scraping is horrible at the best of times but, as anyone who's ever viewed source on one of their pages must be thinking, scraping Google markup must be hell-on-earth.
- userbinator 10y agoIf you have an HTML parser, it's just a matter of selecting the right DOM nodes.
- oneeyedpigeon 10y agoYes, but if your markup is spaghetti, "selecting the right DOM nodes" becomes a lot more difficult.
- audiohihack 10y agoDef not that easy with Google. It is not uncommon to see 20+ serp variations on a given day if you are crawling at high volume, changing user agents, etc. The whole process is fairly terrible to parse consistently.
- Ixiaus 10y agoIt was indeed pretty rough it wouldn't surprise me if Google moves to js generated dom elements to combat rank trackers, at the time it was fine because they want to service non-js browsers but that might change. Parsing it wasn't hard but it wasn't fun...
- slig 10y agoWouldn't it be easier if they generated the dom elements via JS? That would imply that they're getting a JSON or something like, parsing it and creating the DOM.
- Ixiaus 10y agoNo because then you'd have to use a headless browser that can execute js. That increases time and cost when scraping, though it wouldn't surprise me if it ends up going that way.