4 ms·
I'm far from being an AI enthusiast as anyone can be, but this issue has nothing to do with AI specifically. It's just that some greedy companies are writing in
by rnhmjoj 1y ago
I'm far from being an AI enthusiast as anyone can be, but this issue has nothing to do with AI specifically. It's just that some greedy companies are writing incredibly shitty crawlers that don't follow any of the enstablished conventions (respecting robots.txt, using a proper UA string, rate limiting, whatever). This situation could have easily happened earlier than the AI boom, for different reasons.
- mostlysimilar 1y agoBut it didn't, and it's happening now, because of AI.
- kjkjadksj 1y agoPeople have been complaining about these crawlers for years well before AI
- PaulDavisThe1st 1y agoThe issue is 1 to 4 orders of magnitude worse than it was just a couple of years ago. This is not "crawlers suck". This is "crawlers are overwhelming us and almost impossible to fully block". It really isn't the same thing.
- tadfisher 1y agoTragedy of the commons. Before, it was cryptominers eating up all free sources of compute [0]. Now it's AI crawlers eating up all available bandwidth and server resources [1]. Reading SourceHut's struggles against the Once-lers of the world makes me want to introduce a new application layer protocol where consumers pay for abusing shared resources. Which sucks, because the Internet should remain free. [0]: https://drewdevault.com/2021/04/26/Cryptocurrency-is-a-disaster.html https://drewdevault.com/2021/04/26/Cryptocurrency-is-a-disas... [1]: https://drewdevault.com/2025/03/17/2025-03-17-Stop-externalizing-your-costs-on-me.html https://drewdevault.com/2025/03/17/2025-03-17-Stop-externali...
- PaulDavisThe1st 1y ago> Tragedy of the commons. No, because there is no such thing, at least not as understood by Garrett Hardin, who put forward the phrase. Commons fail when selfish, greedy people subvert or destroy the governance structures that help control them. If those governance structures exist (and they do for all historical commons) and continue to exist, the commons suffers no tragedy. This recent slide deck talks about Ostrom's ideas on this, which even Hardin eventually conceded were correct, and that his diagnosis of a "tragedy of the commons" does not actually describe the historical processes by which commons are abused. https://dougwebb.site/slides/commons https://dougwebb.site/slides/commons That said ... arguably there is a problem here with a "commons" that does in fact lack any real governance structure.
- erlend_sh 1y agoNo idea why this is getting downvoted; this is a very important correction since the “tragedy of the commons” meme is based on a flawed premise that needs to be amended.
- p3rls 1y agoi am getting almost 500,000 ai scraper requests a day according to cloudflare's ai audit. google requests the same pages 10+ times each an hour. it was never this bad before.
- majkinetor 1y agoObaying robots.txt can not be enforced. Even if one country makes laws about it, another one will have 0 fucks to give.
- spinningslate 1y agoIt was never intended to be "enforced": > The standard, developed in 1994, relies on voluntary compliance [0] It was conceived in a world with an expectation of collectively respectful behaviour: specifically that search crawlers could swamp "average Joe's" site but shouldn't. We're in a different world now but companies still have a choice. Some do still respect it... and then there's Meta, OpenAI and such. Communities only work when people are willing to respect community rules, not have compliance imposed on them. It then becomes an arms race: a reasonable response from average Joe is "well, OK, I'll allows anyone but [Meta|OpenAI|...] to access my site. Fine in theory, dificult in practice: 1. Block IP addresses for the offending bots --> bots run from obfuscated addresses 2. Block the bot user agent --> bots lie about UA. ...and so on. [0]: https://en.wikipedia.org/wiki/Robots.txt https://en.wikipedia.org/wiki/Robots.txt
- majkinetor 1y agoThanks for the info. However people seem to think that robots.txt will protect them while it was created for another world as you nicelly stated. I guess Nepenthes like tools will be more common in the future, now that tragedy of commons entered digital domain.
- Fomite 1y agoI'd argue it's part of the baked in, fundamental disrespect AI firms have for literally everyone else.
- 1vuio0pswjnm7 1y ago"It's just that some greedy companies are writing incredibly shitty crawlers that don't follow any of the enstablished [sic] conventions (respecting robots.txt, using proper UA string, rate limiting, whatever)." How does "proper UA string" solve this "blowing up websites" problem The only thing that matters with respect to the "blowing up websites" problem is rate-limiting, i.e., behaviour "Shitty crawlers" are a nuisance because of their behaviour, i.e., request rate, not because of whatever UA string they send; the behaviour is what is "shitty" not the UA string. The two are not necessarily correlated and any heuristic that naively assumes so is inviting failure "Spoofed" UA strings have been facilitated and expected since the earliest web browsers For example, https://raw.githubusercontent.com/alandipert/ncsa-mosaic/master/mosaic-spoof-agents https://raw.githubusercontent.com/alandipert/ncsa-mosaic/mas... To borrow the parent's phrasing, the "blowing up websites" problem has nothing to do with UA string specifically It may have something to do with website operator reluctance to set up rate-limiting though; this despite widespread impelementation of "web APIs" that use rate-limiting NB. I'm not suggesting rate-limiting is a silver bullet. I'm suggesting that without rate-limiting, UA string as a means of addressing the "blowing up websites" problem is inviting failure
- AbortedLaunch 1y agoSome of these crawlers appear to be designed to avoid rate limiting based on IP. I regularly see millions of unique ips doing strange requests, each just one or at most a few per day. When a response contains a unique redirect I often see a geographically distinct address fetching the destination.
- 1vuio0pswjnm7 1y ago"I regularly see millions of unique ips doing strange requests, each just one or at most a few per day." How would UA string help For example, a crawler making "strange" requests can send _any_ UA string, and a crawler doing "normal" requests can also send _any_ UA string. The "doing requests" is what I refer to as "behaviour" A website operator might think "Crawlers making strange requests send UA string X but not Y" Let's assume the "strange" requests cause a "website load" problem^1 Then a crawler, or any www user, makes a "normal" request and sends UA string X; the operator blocks or redirects the request, unnecessarily Then a crawler makes "strange" request and sends UA string Y; the operator allows the request and the website "blows up" What matters for the "blowing up websites" problem^1 is behaviour, not UA string 1. The article's title calls it the "blowing up websites" problem, but the article text calls it a problem with "website load". As always the details are missing. For example, what is the "load" at issue. Is it TCP connections or HTTP requests. What number of simultaneous connections and/or requests per second is acceptable, what number is not unacceptable. Again, behaviour is the issue, not UA string The acceptable numbers need to be published; for example, see documentation for "web APIs"
- sznio 1y agoI strongly believe that AI companies are running a DDOS attack on the open web. Making websites go down aligns with their intetests: it removes training data that competitors could use, and it removes sources for humans to browse, making us even more reliant on chatbots to find anything. If it was crap coding, then the bots wouldn't have so many mechanisms to circumvent blocks. Once you block the OpenAI IP ranges, they start using residential proxies. Once you block their UA strings, they start impersonating other crawlers or browsers.