2 ms·
Not only that, even commoncrawl had issues (about a year ago) where AWS couldn't keep up with the demand for downloading the WARCs. As someone who written a lo
by winddude 2y ago
Not only that, even commoncrawl had issues (about a year ago) where AWS couldn't keep up with the demand for downloading the WARCs.
As someone who written a lot of crawling infrastructure and managed large scale crawling operations, respectful crawling is important.
That being said it always seems like google has had a massively unfair advantage for crawling not only with budget but with brandname, and perceived value. It sometimes felt hard to reach out to websites and ask them to allow our crawlers, and grey tactics were often used. And I'm always for a more open internet.
I think regular releases of content in a compressed format would go a long way, but there would always be a race for having the freshest content. What might be better is offering the content in machine format, XML or JSON or even SOAP. Which is usually better for what the sites crawling want to achieve, cheaper for you to serve, and cheaper and less resource intensive compared to crawling. (Have them "cache" locally by enforcing rate limiting and signup)
- apantel 2y ago> That being said it always seems like google has had a massively unfair advantage for crawling not only with budget but with brandname, and perceived value. VCs and other startup culture evangelists are always challenging founders to figure out what their ‘unfair advantage’ is. That’s the name of the game.