4 ms·
It is nice that the AI crawler bots honestly fill out the `User-Agent` header, I'm shocked that they were the source of that much traffic though. 99% of all web
by Proofread0592 1y ago
It is nice that the AI crawler bots honestly fill out the `User-Agent` header, I'm shocked that they were the source of that much traffic though. 99% of all websites do not change often enough to warrant this much traffic, let alone a dev blog.
- grishka 1y agoThey also respect robots.txt. However, I've also seen reports that after getting blocked one way or another, they start crawling with browser user-agents from residential IPs. But it might also be someone else misrepresenting their crawlers as OpenAI/Amazon/Facebook/whatever to begin with.
- rovr138 1y agoWe ended up writing similar rules to the article. It was just based on frequency. While we were rate limiting bots based on UA, we ended up also having to apply wider rules because traffic started spiking from other places. I can't say if it's the traffic shifting, but there's definitely a big amount of automated traffic not identifying itself properly. If you look at all your web properties, look at historic traffic to calculate <hits per IP> in <time period>. Then look at the new data and see how it's shifting. You should be able to identify the real traffic and the automated very quickly.
- deleted 1y ago[deleted]
- cratermoon 1y ago> They also respect robots.txt All the reports I've heard from organizations dealing with AI crawler bots say they are not honest about their user agent and do not respect robots.txt "It's futile to block AI crawler bots because they lie, change their user agent, use residential IP addresses as proxies, and more." https://xeiaso.net/notes/2025/amazon-crawler/ https://xeiaso.net/notes/2025/amazon-crawler/
- eesmith 1y agoFurther info along the same lines at https://drewdevault.com/2025/03/17/2025-03-17-Stop-externalizing-your-costs-on-me.html https://drewdevault.com/2025/03/17/2025-03-17-Stop-externali... > If you think these [AI] crawlers respect robots.txt then you are several assumptions of good faith removed from reality. These bots crawl everything they can find, robots.txt be damned, including expensive endpoints like git blame, every page of every git log, and every commit in every repo, and they do so using random User-Agents that overlap with end-users and come from tens of thousands of IP addresses – mostly residential, in unrelated subnets, each one making no more than one HTTP request over any time period we tried to measure – actively and maliciously adapting and blending in with end-user traffic and avoiding attempts to characterize their behavior or block their traffic. Sourcehut (the site described) used Anubis before swithing "to go-away, which is more configurable and allows us to reduce the user impact of Anubis (e.g. by offering challenges that don’t require JavaScript, or support text-mode browsers better)." https://sourcehut.org/blog/2025-05-29-whats-cooking-q2/ https://sourcehut.org/blog/2025-05-29-whats-cooking-q2/
- immibis 1y agoHowever, there's no evidence those bots are really OpenAI et al.
- grishka 1y agoSpeaking from my own experience, which is admittedly limited, but still — I had AI bots crawling my fediverse server, I added them to my robots.txt as "Disallow: *", they stopped. As I said, it might be someone else entirely using OpenAI/Amazon/Meta/etc user agents to hide their real identity while ignoring robots.txt. What's to stop them? People blame those companies anyway.
- mkfs 1y ago> "It's futile to block AI crawler bots because they lie, change their user agent, use residential IP addresses as proxies, and more." https://xeiaso.net/notes/2025/amazon-crawler/ https://xeiaso.net/notes/2025/amazon-crawler/ If this isn't the case, then the bot detection systems big sites are using must be pretty bad, because I do almost all of my browsing on a desktop originating from residential ASN IP address, and I routinely run up against CAPTCHAs. E.g., any Stack Exchange site on first visit, and even Amazon. What reason would there be for this, unless these crawlers are laundering their traffic through residential IPs?