4 ms·
It's weird that you're dictating rules around scraping not outlined in Github's TOS. The article actually mostly talks about how they scaled working with the a
by cglee 10y ago
It's weird that you're dictating rules around scraping not outlined in Github's TOS.
The article actually mostly talks about how they scaled working with the accumulated data, and only a little about scaling the scraper.
Just like the interesting part about Google is how they process and index the data and not their crawler, the most interesting part of this article actually not the scraping, but how they handled all the data and processed it.
Edit: I found Github's policy on scraping. I hope the link brings some closure to this concern: https://help.github.com/articles/github-terms-of-service-draft/#5-scraping https://help.github.com/articles/github-terms-of-service-dra...