3 ms·
Writing distributed and highly-concurrent software in Erlang to scrape Google's search results, with localization, tens of millions of times per-day. I don't w
by ixmatus 12y ago
Writing distributed and highly-concurrent software in Erlang to scrape Google's search results, with localization, tens of millions of times per-day.
I don't work on that stuff anymore, but it definitely challenged my problem solving abilities to:
A) Learn Erlang.
B) Learn how to write solid distributed software in Erlang.
C) Figure out how to work around Google's temp-banning policies using IP balancing, captcha solving to produce cookies to balance connections on, and also how to localize the searches for geographic accuracy.
I don't like working on projects that are actively pitting me against someone though, so I'm happy to not be working on that. I now write scalable software for my energy-focused startup, we receive energy data from homes in near-real time, which has its own challenges.
- sireat 12y agoSounds very challenging, yet also very black-hattish on a massive scale not that I have much love for Google. That is the problem that many of the most interesting projects have some moral ambiguities(military, financial, etc, etc). While one is getting paid, it is very easy to justify or not even think about where the money comes from(que Sinclair quote).
- ixmatus 12y agoNot quite black-hat; I consider tiered link building to be more blackhat. Scraping the search results was about getting data on where items were positioned in the index, not pumping spam into the search results. Although to be clear, I did help do stuff like that for a while too but it never felt good and we vastly preferred the rank tracking product to link building services. Also, Google's getting extremely good at combating search spam.
- dewitt 12y ago> I don't like working on projects that are actively pitting me against someone though. I'm surprised you didn't also didn't mention the ethical problems with trying to take something (search results) without permission. Your new project sounds awesome, though. : )
- patio11 12y agoGoogle doesn't exactly tie itself in knots that they built a few hundred billion dollars on top of "We're going to crawl the entire Internet and datamine it for our own purposes. Nobody will agree to this, so we won't ask them. Instead, we will offer easy ways to opt out after having achieved hegemonic control of Internet navigation."
- ixmatus 12y agoMeh. I didn't have any qualms with that, per-se. As others have pointed out I think there are ethical concerns with what Google is doing itself. Also it's indicative of something if people are building businesses to scrape something they could offer as a pay-for API that many people would gladly / happily pay lots of money for. The arguments for NOT doing that are ridiculous because people will get the data regardless.