3 ms·
IA has not honored robots.txt for the better part of a decade now. https://blog.archive.org/2017/04/17/robots-txt-meant-for-search-engines-dont-work-well-for-w
by walski 1y ago
IA has not honored robots.txt for the better part of a decade now.
https://blog.archive.org/2017/04/17/robots-txt-meant-for-search-engines-dont-work-well-for-web-archives/ https://blog.archive.org/2017/04/17/robots-txt-meant-for-sea...
- lxgr 1y agoAre you sure? The article (from 2017) you've linked only mentions "U.S. government and military web sites", and their wayback machine FAQ still mentions that robots.txt "might" prevent crawling: https://help.archive.org/help/using-the-wayback-machine/ https://help.archive.org/help/using-the-wayback-machine/