3 ms·
Correct. Example snippet from the nytimes.com robots.txt: User-agent: archive.org_bot Disallow: /
by cmeacham98 5mo ago
Correct. Example snippet from the nytimes.com robots.txt:
User-agent: archive.org_bot
Disallow: /
- joecool1029 5mo agoWhich they don’t respect. I’ve had it for my blog for years and they still added it to wayback machine, see my last comment for their official announcement of the ignore robots.txt policy, it is not new.
- socalgal2 5mo agorobots.txt means they shouldn't auto-scan your site. Any user though can go to the wayback machine and type in a URL and the wayback machine will read that URL. That was the intent of robots.txt (don't scan) not (don't read period). It's spelled out in the spec for robots.txt
- keane 5mo agoThe <meta name="robots"> tag and robots.txt serve different roles: robots.txt controls crawling, while the robots meta tag influences indexing and other behavior. https://developer.mozilla.org/en-US/docs/Web/HTML/Reference/Elements/meta/name/robots https://developer.mozilla.org/en-US/docs/Web/HTML/Reference/... I wonder how archive.org_bot behaves when <meta name="robots" content="noindex, noarchive, nocache" /> is present.
- socalgal2 5mo agoThe person above those is complaining about entries in their logs from bots. A robot can't read a tag without first reading the document. So sure, if they're a good bot they might not store the results but the server's logs will still show the bot's GET request.
- ninjagoo 5mo ago> I’ve had it for my blog for years Just out of curiosity, why don't you want your public blog archived? not questioning, just trying to understand the logic/motivations? Also, I think you're being unfairly downvoted.
- mjmas 5mo agoIs there a difference between that and User-agent: ia_archiver ?
- keane 5mo agoThat was operated by Alexa Internet and the results powered the Internet Archive (same founder) presumably until their acquisition: https://web.archive.org/web/20140528103516/https://alexa.zendesk.com/hc/en-us/articles/200450194-Alexa-s-Web-and-Site-Audit-Crawlers https://web.archive.org/web/20140528103516/https://alexa.zen... https://en.wikipedia.org/wiki/Alexa_Internet https://en.wikipedia.org/wiki/Alexa_Internet