6 ms·
What do you mean by "barely" respecting robots.txt? Wouldn't that be more binary? Are they respecting some directives and ignoring others?
by gundmc 2y ago
What do you mean by "barely" respecting robots.txt? Wouldn't that be more binary? Are they respecting some directives and ignoring others?
- unsnap_biceps 2y agoI believe that a number of AI bots only respect robot.txt entries that explicitly define their static user agent name. They ignore wildcards in user agents. That counts as barely imho. I found this out after OpenAI was decimating my site and ignoring the wildcard deny all. I had to add entires specifically for their three bots to get them to stop.
- noman-land 2y agoThis is highly annoying and rude. Is there a complete list of all known bots and crawlers?
- jsheard 2y agohttps://darkvisitors.com/agents https://darkvisitors.com/agents https://github.com/ai-robots-txt/ai.robots.txt https://github.com/ai-robots-txt/ai.robots.txt
- joecool1029 2y agoEven some non-profit ignore it now, Internet Archive stopped respecting it years ago: https://blog.archive.org/2017/04/17/robots-txt-meant-for-search-engines-dont-work-well-for-web-archives/ https://blog.archive.org/2017/04/17/robots-txt-meant-for-sea...
- SR2Z 2y agoIA actually has technical and moral reasons to ignore robots.txt. Namely, they want to circumvent this stuff because their goal is to archive EVERYTHING.
- amarcheschi 2y agoI also don't think they hit servers repeatedly so much
- prinny_ 2y agoIsn’t this a weak argument? OpenAI could also say their goal is to learn everything, feed it to AI, advance humanity etc etc.
- AnonC 2y agoAs I recall, this is outdated information. Internet Archive does respect robots.txt and will remove a site from its archive based on robots.txt. I have done this a few years after your linked blog post to get an inconsequential site removed from archive.org.
- dredmorbius 2y agoThe most recent notice IA have blogged was in 2017, and there's no indication that the service has reversed course on robots.txt since. <https://blog.archive.org/?s=robots.txt https://blog.archive.org/?s=robots.txt>
- deleted 2y ago[deleted]
- LukeShu 2y agoAmazonbot doesn't respect the `Crawl-Delay` directive. To be fair, Crawl-Delay is non-standard, but it is claimed to be respected by the other 3 most aggressive crawlers I see. And how often does it check robots.txt? ClaudeBot will make hundreds of thousands of requests before it re-checks robots.txt to see that you asked it to please stop DDoSing you.
- mariusor 2y agoOne would think they'd at least respect the cache-control directives. Those have been in the web standards since forever.
- Animats 2y agoHere's Google, complaining of problems with pages they want to index but I blocked with robots.txt. New reason preventing your pages from being indexed Search Console has identified that some pages on your site are not being indexed due to the following new reason: Indexed, though blocked by robots.txt If this reason is not intentional, we recommend that you fix it in order to get affected pages indexed and appearing on Google. Open indexing report Message type: [WNC-20237597]
- smarnach 2y agoThey are not complaining. You configured Google Search Console to notify you about problems that affect the search ranking of your site, and that's what they do. if you don't want to receive these messages, turn them off in Google Search Console.