3 ms·
Lawsuit nastiness aside, there's an interesting and important legal-technical question that this exposes: how should websites specify acceptable uses of crawled
by randomwalker 16y ago
Lawsuit nastiness aside, there's an interesting and important legal-technical question that this exposes: how should websites specify acceptable uses of crawled data and other fine-grained restrictions in a machine-readable form.
Motivated by this incident, I got together with Pete (the author/victim) to write a piece on "The Need to Reboot Robots.txt" [1] but it went nowhere.
Any suggestions on how to give our proposal legs would be much appreciated.
[1] http://33bits.org/2010/12/05/web-crawlers-privacy-reboot-robots-txt/ http://33bits.org/2010/12/05/web-crawlers-privacy-reboot-rob...
- deleted 16y ago[deleted]
- blauwbilgorgel 16y agoYou can set the SyndicationRight directive for OpenSearch. "Contains a value that indicates the degree to which the search results provided by this search engine can be queried, displayed, and redistributed." The default is "open" meaning: - The search client may request search results. - The search client may display the search results to end users. - The search client may send the search results to other search clients. http://www.opensearch.org/Specifications/OpenSearch/1.1#The_.22SyndicationRight.22_element http://www.opensearch.org/Specifications/OpenSearch/1.1#The_... That would give you more fine-grained control over what search agents do with your data. I don't know how broad the support and adherence is to the OpenSearch spec (IMDB uses it).