5 ms·
You call it extortion of the AI companies, but isn’t stealing/crawling/hammering a site to scrape their content to resell just as nefarious? I would say Cloudfl
by james2doyle 10mo ago
You call it extortion of the AI companies, but isn’t stealing/crawling/hammering a site to scrape their content to resell just as nefarious? I would say Cloudflare is giving these site owners an option to protect their content and as a byproduct, reduce their own costs of subsidizing their thieves. They can choose to turn off the crawl protection. If they aren't, that tells you that they want it, doesn’t it?
- cpncrunch 10mo ago>You call it extortion of the AI companies, but isn’t stealing/crawling/hammering a site to scrape their content to resell just as nefarious? You can easily block ChatGPT and most other AI scrapers if you want: https://habeasdata.neocities.org/ai-bots https://habeasdata.neocities.org/ai-bots
- james2doyle 10mo agoThis is just using robots.txt and asking "pretty please, don’t scrape me". Here is an article (from TODAY) about the case where Perplexity is being accused of ignoring robots.txt: https://www.theverge.com/news/839006/new-york-times-perplexity-lawsuit-copyright https://www.theverge.com/news/839006/new-york-times-perplexi... If you think a robots.txt is the answer to stopping the billion-dollar AI machine from scraping you, I don’t know what to say.
- cpncrunch 10mo agoYes, I was referring to legitimate companies, and Perplexity doesn't seem to be one of those.
- albedoa 10mo agoOh for sure. When he wrote of the AI companies that are "stealing/crawling/hammering", you thought he meant the legitimate ones that do honor robots.txt. That makes sense.
- cpncrunch 10mo agoActually, it looks like all the major ones do honour robots.txt including perplexity. They seemingly get around it using google serps, so theyre not actually crawling or hammering the site servers (or even cloudflare). https://www.ailawandpolicy.com/2025/10/anti-circumvention-reddits-case-against-perplexity/ https://www.ailawandpolicy.com/2025/10/anti-circumvention-re...
- Aeolun 10mo agoIf someone has a robots.txt, and I want to request their page, but I want to do that in an automated way, should I open the browser to do it instead of issue a curl request? How about if I am going to ask claude to fetch the page for me?
- kentm 10mo agoRespect the robots.txt and don’t do it?
- jacobgkau 10mo agoI'm guessing you don't manage any production web servers? robots.txt isn't even respected by all of the American companies. Chinese ones (which often also use what are essentially botnets in Latin American and the rest of the world to evade detection) certainly don't care about anything short of dropping their packets.
- dingnuts 10mo ago[dead]
- cpncrunch 10mo agoI have been managing production commercial web servers for 28 years. Yes, there are various bots, and some of the large US companies such as Perplexity do indeed seem to be ignoring robots.txt. Is that a problem? It's certainly not a problem with cpu or network bandwidth (it's very minimal). Yes, it may be an issue if you are concerned with scraping (which I'm not). Cloudflare's "solution" is a much bigger problem that affects me multiple times daily (as a user of sites that use it), and those sites don't seem to need protection against scraping.
- filleduchaos 10mo agoIt is rather disingenuous to backpedal from "you can easily block them" to "is that a problem? who even cares" when someone points out that you cannot in fact easily block them.
- cpncrunch 10mo agoI was referring to legitimate ones, which you can easily block. Obviously there are scammy ones as well, and yes it is an issue, but for most sites I would say the cloudflare cure is worse than the problem it's trying to cure.
- oasisbob 10mo agoNo true scotsman needs Cloudflare, as any true scotsman can block AI bots themselves is not a strong argument.
- mplewis 10mo agoNo you cannot! I blocked all of the user agents on a community wiki I run, and the traffic came back hours later masquerading as Firefox and Chrome. They just fucking lie to you and continue vacuuming your CPU.
- cpncrunch 10mo agoThere shouldn't be any noticeable hit on your cpu from bots from a site like that. Are you sure it's not a DDoS? Obviously it depends on the bot, and you can't block the scammy ones. I was really just referring to the major legitimate companies (which might not include Perplexity).
- literalAardvark 10mo agoThere is a noticeable hit, there's also a noticeable cost, and it's not a ddos. Not all sites can have full caching, we've tried.
- cpncrunch 10mo agoI was referring to the community wiki.
- deleted 10mo ago[deleted]
- chrneu 10mo agothis is the equivalent of asking people not to speed on your street.
- literalAardvark 10mo agoTell me you don't run a site without telling me you don't run a site
- cpncrunch 10mo agoTell me you make incorrect assumptions without specifically saying so. (Yes, you're incorrect).
- Sohcahtoa82 10mo agoHow are you this naive? Do you really think scrapers give a damn about your robots.txt?
- cpncrunch 10mo agoThe legitimate ones do, which is what I was referring to. Obviously there are bastard ones as well.