2 ms·
Although strictly speaking you're right, it breaks the intent of robots.txt as the users' clickstream is fed into the server-side indexing engine alongside robo
by motti 16y ago
Although strictly speaking you're right, it breaks the intent of robots.txt as the users' clickstream is fed into the server-side indexing engine alongside robot crawled data.
Arguing that according to the strict letter of the spec, robots.txt only has to be obeyed by true crawlers does hold water as strictly speaking Bing could disregard robots.txt altogether - it's not certainly not legally enforceable. The intent of robots.txt is clear and Bing should be trying to obey it wherever possible.
In this case it appears the keyword was associated because it was typed into the Google search box, and from the there to the faked destination. IMHO when the clickstream was analysed it should have disregarded clickstreams that pass a robots.txt-excluded page as these could establish associations that were not supposed to revealed by crawlers.
We aught to hold Bing and any search engine to the highest standard when talking about crawling etiquette.
Also, see http://news.ycombinator.com/item?id=2169817 http://news.ycombinator.com/item?id=2169817
- Dylan16807 16y agoIf I'm google and I block /search I want no company to collect the data. If I'm a site that blocks /dynamic because I want to make sure my stats are accurate, not factoring bots, and I want to keep the load down on slow-generating pages, then I'm perfectly happy with the data being collected. Is it the intent of robots.txt itself to block the clickstream data, or is it just google's intent?