4 ms·
"Some data is simply not worth indexing, and not worth serving up to bots." The broader points being made in the article are that the value of information is d
by panza 15y ago
"Some data is simply not worth indexing, and not worth serving up to bots."
The broader points being made in the article are that the value of information is determined by the visitor, and that the burden of keeping the site up and running should fall on the host.
By that reasoning, the presence of a ROBOTS.TXT signifies "damage" or "temporary madness", and should therefore be ignored.
Also, keep in mind where this article came from - the Archive Team are bloody-minded about preserving information.
- drivingmenuts 15y agoOK, wait. If the data has value to the user, shouldn't the user be paying the host for the cost of making that data available (or perhaps considerably more, if it has a lot of value to the user)?
- sbierwagen 15y agoExcellent. How much are you paying ycombinator for your account? If the answer is "nothing", then I guess you've just told me that all your comments are useless, and I should ignore them.
- sbierwagen 15y agoI'm committing the classic mistake of talking about the karma system here, but what, precisely, is wrong with my statement? Drivingmenuts says the user should pay for hosting data. HN is hosting his data, but he doesn't pay for it. Is there some error in my logic, here, some flaw in my conclusion?
- smosher 15y agothe value of information is determined by the visitor That may be so in a sense, unless the visitor has no means of making the call -- or does so poorly. I promise you that every robots.txt ignoring bot makes terrible judgments in that regard. (I'm not trying to overstate anything here; those that abide it make, on average, slightly less terrible judgments.) But it's not just the visitor who gets to make judgment calls. The value of serving those pages is something the host can decide. Belligerent drunks who abuse the staff aren't allowed in the coffee shop, and if they do it often enough they're not allowed back in when they sober up either.
- chalst 15y agothe Archive Team are bloody-minded about preserving information They won't do a good job of that if they encourage too many people to make IP blocks of their machines.
- sbierwagen 15y agoAT, the organization, doesn't own any machines. Operations are all done by members, the majority of whom use consumer ISPs.
- chalst 15y agoDo all of them ignore robots.txt?
- sbierwagen 15y agoWell, when I iterated through the everything2.com namespace and downloaded the majority of their content[1], I respected their robots.txt. After I did that, they changed it to: User-agent: * Disallow: / Does that count? 1: http://bbot.org/blog/archives/2011/01/17/more_fun_with_wget/ http://bbot.org/blog/archives/2011/01/17/more_fun_with_wget/