9 ms·
Yep -- our story here: https://about.readthedocs.com/blog/2024/07/ai-crawlers-abuse/ https://about.readthedocs.com/blog/2024/07/ai-crawlers-abuse... (quoted in
by ericholscher 2y ago
Yep -- our story here: https://about.readthedocs.com/blog/2024/07/ai-crawlers-abuse/ https://about.readthedocs.com/blog/2024/07/ai-crawlers-abuse... (quoted in the OP) -- everyone I know has a similar story who is running large internet infrastructure -- this post does a great job of rounding a bunch of them up in 1 place.
I called it when I wrote it, they are just burning their goodwill to the ground.
I will note that one of the main startups in the space worked with us directly, refunded our costs, and fixed the bug in their crawler. Facebook never replied to our emails, the link in their User Agent led to a 404 -- an engineer at the company saw our post and reached out, giving me the right email -- which I then emailed 3x and never got a reply.
- pjc50 2y ago> just burning their goodwill to the ground AI firms seem to be leading from a position that goodwill is irrelevant: a $100bn pile of capital, like an 800lb gorilla, does what it wants. AI will be incorporated into all products whether you like it or not; it will absorb all data whether you like it or not.
- anthk 2y agoAI tarpits && lim (human curated contant/mediocre AI answers -> 0) = AI's crumbling into dust by themselves.
- UncleMeat 2y agoYep. And it is much more far reaching than that. Look at the primary economic claim offered by AI companies: to end the need for a substantial portion of all jobs on the planet. The entire vision is to remake the entire world into one where the owners of these companies own everything and are completely unconstrained. All intellectual property belongs to them. All labor belongs to them. Why would they need good will when they own everything? "Why should we care about open source maintainers" is just a microcosm of the much larger "why should we care about literally anybody" mindset.
- chii 2y ago> remake the entire world into one where the owners of these companies own everything and are completely unconstrained how has this been any different from the past 10,000 years of human conquest and domination?
- nemomarx 2y agoin the past, you had to give some of your spoils to those who did the conquering for you, and laborers after that. if you can automate and replace all work, including maintening the robots that do that and training them, you no longer need to share anything.
- lithocarpus 2y agoIn my view it's the same thing, same trajectory -- with more power in the hands of fewer people further along the trajectory. It can be better or worse depending on what those with power choose to do. Probably worse. There has been conquest and domination for a long time, but ordinary people have also lived in relative peace gathering and growing food in large parts of the world in the past, some for entire generations. But now the world is rapidly becoming unable to support much of that as abundance and carrying capacity are deleted through human activity. And eventually the robot armies controlled by a few people will probably extract and hoard everything that's left. Hopefully in some corners some people and animals can survive, probably by being seen as useful to the owners.
- neutronicus 2y agoOn the bright side, armies of robot slaves give us an off-ramp from the unsustainable pyramid scheme of population growth. Be fruitful, and multiply, so that you may enjoy a comfortable middle age and senescence exploiting the shit out of numerous naive 25-year-olds! If it's robots, we can ramp down the population of both humans and robots until the planet can once again easily provide abundance.
- lithocarpus 2y ago
- davidmurdoch 2y agoWe, the people, might need to come up with a few proverbial tranquilizer guns here soon
- ferguess_k 2y agoThat's pretty much what our future would look like -- you are irrelevant. Well I mean we are already pretty much irrelevant nowadays, but the more so in the "progressive" future of AI.
- huijzer 2y agoI think the logic is more like “we have to do everything we can to win or we will disappear”. Capitalism is ruthless and the big techs finally have some serious competition, namely: each other as well as new entrants. Like why else can we just spam these AI endpoints and pay $0.07 at the end of the month? There is some incredible competition going on. And so far everyone except big tech is the winner so that’s nice.
- Sharlin 2y agoMaxim 1: "Pillage, then burn."
- Coffeewine 2y agoAnother Schlock Mercenary fan? Or does this adage have many adherents?
- datadrivenangel 2y agoThe adage predates the longest continuous webcomic, but not as a maxim.
- Sharlin 2y agoYep, a fan I am.
- yubblegum 2y agoThey are also gutting the profession of software engineering. It's a clever scam actually: to develop software a company will need to pay utility fees to A"I" companies and since their products are error prone voila use more A"I" tools to correct the errors of the other tools. Meanwhile software knowledge will atrophy and soon ala WALE we'll have software "developers" with 'soft bones' floating around on conveyed seats slurping 'sugar water' and getting fat and not knowing even how to tie their software shoelaces.
- kordlessagain 2y ago> AI will be incorporated into all products whether you like it or not AI will be incorporated into the government, whether you like it or not. FTFY!
- slowmovintarget 2y ago"... you have the lawyers clean it all up later." - Eric Schmidt
- b112 2y agoYes, like the Pixel camera app, which mangles photos with AI processing, and users complain that it won't let people take pics. One issue was a pic with text in it, like a store sign. Users were complaining that it kept asking for better focus on the text in the background, before allowing a photo. Alpha quality junk. Which is what AI is, really.
- asveikau 2y agoRules and laws are for other people. A lot of people reading this comment having mistaken "fake it til you make it" or "better to not ask permission" for good life advice are responsible for perpetrating these attitudes, which are fundamentally narcissistic.
- speerer 2y agohttps://arstechnica.com/ai/2025/03/devs-say-ai-crawlers-dominate-traffic-forcing-blocks-on-entire-countries/ https://arstechnica.com/ai/2025/03/devs-say-ai-crawlers-domi... links to this comment.
- ferguess_k 2y agoMaybe just feed them dynamically generated garbage information? More fun than no information.
- InfamousRece 2y agoIt does not even have to be dynamically generated. Just pre-generate a few thousand static pages of AI slop and serve that. Probably cheaper than dynamic generation.
- gnz11 2y agoOP’s linked blog post mentioned they got hit with a large spike in bandwidth charges. Sending them garbage information costs money.
- ferguess_k 2y agoYeah you have a point, hmmm, wish there were a way to somehow generate those garbages with minimum bandwidth. Something like, I can send you a very compressed 256 bytes of data which expands to something like 1 mega bytes.
- kevindamm 2y agothere is -- but instead of garbage expanding data, add in several delays within the response so that the data takes extraordinarily long Depending on the number of simultaneous requesting connections, you may be able to do this without a significant change to your infrastructure. There are ways to do it that don't exhaust your number of (IP, port) available too, if that is an issue. Then the hard part is deciding which connections to slow, but you can start with a proportional delay based on the number of bytes per source IP block or do it based on certain user agents. Might turn into a small arms race but it's a start.
- madeforhnyo 2y agoGood ol' zip bomb https://furry.engineer/@niko/113728467796605323 https://furry.engineer/@niko/113728467796605323
- spenczar5 2y agoThanks for writing about this. Is it clear that this is from crawlers, as opposed to dynamic requests triggered by LLM tools, like Claude Code fetching docs on the fly?
- TuringNYC 2y ago>> which I then emailed 3x and never got a reply. At which point does the crawling cease to be a bug/oversight and constitute a DDOS?
- lgeek 2y ago> One crawler downloaded 73 TB of zipped HTML files in May 2024 [...] This cost us over $5,000 in bandwidth charges I had to do a double take here. I run (mostly using dedicated servers) infrastructure that handles a few hundred TB of traffic per month, and my traffic costs are on the order of $0.50 to $3 per TB (mostly depending on the geographical location). AWS egress costs are just nuts.
- Ray20 2y agoI think uncontrolled price of cloud traffic - is a real fraud and way bigger problem then some AI companies that ignore robot.txt. One time we went over limit on Netlify or something, and they charged over thousand for a couple TB.
- joepie91_ 2y ago> I think uncontrolled price of cloud traffic - is a real fraud Yes, it is. > and way bigger problem then some AI companies that ignore robot.txt. No, it absolutely is not. I think you underestimate just how hard these AI companies hammer services - it is bringing down systems that have weathered significant past traffic spikes with no issues, and the traffic volumes are at the level where literally any other kind of company would've been banned by their upstream for "carrying out DDoS attacks" months ago.
- Ray20 2y ago>I think you underestimate just how hard these AI companies hammer services Yeas, I completely don't understand this and don't understand comparing this with ddos attacks. There's no difference with what search engines are doing, and in some way it's worse? How? It's simply scraping data, what significant problems may it cause? Cache pollution? And thats'it? I mean even when we talking about ignoring robots.txt (which search engines are often doing too) and calling costly endpoints - what is the problem to add to those endpoints some captcha or rate limiters?
- Suppafly 2y ago>which I then emailed 3x and never got a reply. Send a bill to their accounts payable team instead.
- ldoughty 2y agoDetect AI scraper and inject an in-page notice that by continuing they accept your terms of use. Terms of use charges them per page load in some terminology of abuse. Profit... By sending them invoices :-)
- dabockster 2y agoHonestly this is crazy enough to work. Bonus points if both you and the scraping company reside in the same state.
- Freebytes 2y agoAlong with having block lists, perhaps you could add poison to your results that generates random bad code that will not work, and that is only seen by bots (display: none when rendered), and the bots will use it, but a human never would.
- ATechGuy 2y agoWondering if used tried stopping such bots with Captcha?
- m463 2y agoI kind of suspect some of these companies probably have more horsepower and bandwidth in one crawler than a lot of these projects have in their entire infrastructure.