3 ms·
(Other co-founder of Neeva) We wrote an op-ed last month in Fast Company re: the challenges of crawling at scale in today's world: https://www.fastcompany.com/
by vivekraghunatha 4y ago
(Other co-founder of Neeva)
We wrote an op-ed last month in Fast Company re: the challenges of crawling at scale in today's world: https://www.fastcompany.com/90759792/with-google-dominating-search-the-internet-needs-crawl-neutrality https://www.fastcompany.com/90759792/with-google-dominating-...
tldr; there are two big challenges in crawling the web:
* quantity -- how do you crawl the web at O(B) pages per day
* quality -- how do you make sure you are crawling the high quality parts
On the quantity side, the web is a much trickier place to crawl than 10-15 years ago. Even if you build a system that is well behaved, respects rate limits and work w/ webmasters and CDNs to make sure you are treated as a good bot, the following things will bite you:
* some sites only allow googlebot/bingbot via robots
* even when sites allow all search crawlers, you'll see 429s, 503s, crawl delay directives that limit you to very low qps, and other mysterious errors. (most retailers and aggregators)
* working w/ webmasters works only if you manage to get a response from them (good luck)
* many sites require JS rendering (which is 100x more expensive in terms of the number of assets you are crawling)
(OTOH, Cloudflare is a force for good: https://radar.cloudflare.com/verified-bots https://radar.cloudflare.com/verified-bots)
On the quality side, crawl prioritization works best when you have click data, which most small search engines don't have enough of. In the absence of that, good seed sets and high quality inlinks are your friend.
In other news, eng blogs are being written up as we speak. Will post here on HN as they roll out. Also, like marginalia_nu pointed out, we share whatever we've been working on on a weekly basis on our Twitter account.