3 ms·
Think of the potential alternatives. One is everything can be crawled with no way for the site owner to say "no, please don't crawl this." The other major one
by elehack 15y ago
Think of the potential alternatives. One is everything can be crawled with no way for the site owner to say "no, please don't crawl this."
The other major one is to legally bar all crawling without express permission.
The current de facto world - crawling is OK unless robots.txt says otherwise - is pretty nice. If we want that to be a legal defense in court ("You didn't put up a robots.txt, so my indexing was legal, so you can't sue me."), which seems useful, then the necessary flipside of that is that violating robots.txt exposes the crawler to liability. That's a tradeoff I'm perfectly willing to accept to allow the web, and necessary services such as indexers and crawlers, to work while still allowing publishers to have some reasonable control over their content distribution.
I seem to remember one of the writers at Search Engine Land presenting a nice description of the robots.txt request in contract negotiation terms. Something like this:
Archiver: "Are there any limits on what I can archive or index from your site?" (translation: GET /robots.txt)
Site: "Nope." (translation: 404 Not Found)
or
Site: "Yep, here they are." (translation: 200 OK followed by restrictions in robots.txt format)
So, by asking for /robots.txt, the crawler can be construed as asking permission to index, and the response setting up the terms of indexing. That seems like a really useful defense and sane compromise in this age of "indexing so you can drive search users to our content is copyright infringement."
[EDIT: fix formatting]
- _delirium 15y agoI guess the first alternative seems like the more sane one to me, as far as legality goes. Website-crawling policies don't seem like the kind of thing that rises to the level where it's worth involving courts and laws, so I'd leave it to technological mechanisms plus voluntary compliance with non-technological mechanisms. But I suppose I have a pretty high bar for what problems are severe enough to require a government solution. And pragmatically, the vast majority of the non-robots.txt-respecting crawls I see are coming from countries that won't enforce such laws anyway, so enforcing them in western countries seems like a downside (more entanglement between the internet and various countries' national laws) with little upside (won't stop many crawls).