3 ms·
> archive.org's copy goes down. From asking around they do retain the original data, but exclude it from public results following a robots.txt exclusion. It's
by Springtime 12y ago
> archive.org's copy goes down.
From asking around they do retain the original data, but exclude it from public results following a robots.txt exclusion. It's something I do wish they would re-consider though as it limits the purpose.
- chroma 12y agoI hope your claim about them keeping the data is correct, but... If a tree falls in the woods and robots.txt later excludes it, does archive.org make a sound? While the distinction may be useful to someone with access to the data, it doesn't matter to me. From my point of view, the URL goes from available to unavailable. It's indistinguishable from typical link rot. The Internet Archive has kept their robots.txt policy for over a decade now, despite constant requests to change it. I doubt they'll change it any time soon.
- cogburnd02 12y agoThere's a forum here: https://archive.org/post/406632/why-does-the-wayback-machine-pay-attention-to-robotstxt https://archive.org/post/406632/why-does-the-wayback-machine... Maybe if it gets enough attention, and/or the Internet Archive people get enough e-mails about this problem (Wayback Machine obeying robots.txt) then they'll change their mind. https://archive.org/about/bios.php https://archive.org/about/bios.php Here's a link with the director's email: http://brewster.kahle.org/about/ http://brewster.kahle.org/about/