2 ms·
I think the goal is to prevent other robots from accidentally crawling the Google corpus. At the end of the day, robots.txt is a stab at comity. There are many
by dschoon 17y ago
I think the goal is to prevent other robots from accidentally crawling the Google corpus. At the end of the day, robots.txt is a stab at comity. There are many well-known services which use crawlers that fail to obey it (or identify themselves). The file merely announces, "This is all recycled content." -- notice that none of the Google corporate pages are disallowed.
Aside: let's say you've whipped up a spiffy new ranking algorithm, and you just need an index to launch your search engine. What's faster: crawling the web, or crawling Google? I don't think such an entrepreneur would pass on a big speed up just because of a text-file.