5 ms·
https://web.archive.org/web/20150811052336/https://blogs.oracle.com/maryanndavidson/entry/no_you_really_can_t https://web.archive.org/web/20150811052336/https:/
by anglebracket 11y ago
https://web.archive.org/web/20150811052336/https://blogs.oracle.com/maryanndavidson/entry/no_you_really_can_t https://web.archive.org/web/20150811052336/https://blogs.ora...
- opless 11y agoJust in case a robot.txt kills that http://pastebin.com/rcPSyRnR http://pastebin.com/rcPSyRnR
- snsr 11y agoIt's also on seclists.org - http://seclists.org/isn/2015/Aug/4 http://seclists.org/isn/2015/Aug/4
- hughw 11y agoWould archive.org typically honor a robots.txt for a resource it already retrieved? I never understood the intent of a robots.txt to be retroactive.
- mikeash 11y agoApparently yes, it would: https://archive.org/about/exclude.php https://archive.org/about/exclude.php
- X-Istence 11y agoYes, it simply hides the content, it is still kept in their database so if the robots.txt disappears, it pops back from their archive. New pages won't be archived though.
- syncsynchalt 11y agoMy understanding is that sites like archive.org honor robots.txt retroactively not because they are required to, but to best honor the wishes of the content provider.