7 ms·
Cloudflare crawl endpoint
- rvz 7mo agoSelling the cure (DDoS protection) and creating the poison (Authorized AI crawling) against their customers.
- triwats 7mo agothis could be cool to use cloudflare's edge to do some monitoring of endpoints actual content for synthetic monitoring
- jasongill 7mo agoI'm surprised that Cloudflare hasn't started hosting a pre-scraped version of websites that use Cloudflare's proxy - something like https://www.example.com/cdn-cgi/cached-contents.json https://www.example.com/cdn-cgi/cached-contents.json They already have the website content in their cache, so why not just cut out the middle man of scraping services and API's like this and publish it? Obviously there's good reasons NOT to, but I am surprised they haven't started offering it (as an "on-by-default" option, naturally) yet.
- cmsparks 7mo agoThat would prolly work for simple sites, but you still need the dedicated scraping service with a browser to render sites that are more complex (i.e. SPAs)
- michaelmior 7mo ago> I'm surprised that Cloudflare hasn't started hosting a pre-scraped version of websites that use Cloudflare's proxy It's entirely possible that they're doing this under the hood for cases where they can clearly identify the content they have cached is public.
- binarymax 7mo agoBased on the post, it seems likely that they'd just delay per the robots.txt policy no matter what, and do a full browser render of the cached page to get the content. Probably overkill for lots and lots of sites. An HTML fetch + readability is really cheap.
- janalsncm 7mo agoHow would they know the content hasn’t changed without hitting the website?
- coreq 7mo agoThey wouldn't, well there's Etag and alike but it still a round trip on level 7 to the origin. However the pattern generally is to say when the content is good to in the Response headers, and cache on that duration, for an example a bitcoin pricing aggregator might say good for 60 seconds (with disclaimers on page that this isn't market data), whilst My Little Town news might say that an article is good for an hour (to allow Updates) and the homepage is good for 5 minutes to allow breaking news article to not appear too far behind.
- OptionOfT 7mo agoCaching headers? (Which, on Akamai, are by default ignored!)
- cortesoft 7mo agoKeeping track of when content changes is literally the primary function of a CDN.
- csomar 7mo agoIt’s a bit more complicated than that. This is their product Browser Rendering, which runs a real browser that loads the page and executes JavaScript. It’s a bit more involved than a simple curl scraping.
- randomtools 7mo agoSo does that mean it can replace serpapi or similar?
- selcuka 7mo agoNot the same thing, but they have something close (it's not on-by-default, yet) [1]: > Cloudflare's network now supports real-time content conversion at the source, for enabled zones using content negotiation headers. Now when AI systems request pages from any website that uses Cloudflare and has Markdown for Agents enabled, they can express the preference for text/markdown in the request. Our network will automatically and efficiently convert the HTML to markdown, when possible, on the fly. [1] https://blog.cloudflare.com/markdown-for-agents/ https://blog.cloudflare.com/markdown-for-agents/
- cortesoft 7mo agoWell, the conversion process into the JSON representation is going to take CPU, and then you have to store the result, in essence doubling your cache footprint. Doing it on demand still utilizes their cached version, so it saves a trip to the origin, but doesn’t require doubling the cache size. They can still cache the results if the same site is scraped multiple times, but this saves having to cache things that are never going to be requested. Cache footprint management is a huge factor in the cost and performance for a CDN, you want to get the most out of your storage and you want to serve as many pages from cache as possible. I know in my experience working for a CDN, we were doing all sorts of things to try to maximize the hit rate for our cache.. in fact, one of the easiest and most effective techniques for increasing cache hit rate is to do the OPPOSITE of what you are suggesting; instead of pre-caching content, you do ‘second hit caching’, where you only store a copy in the cache if a piece of content is requested a second time. The idea is that a lot of content is requested only once by one user, and then never again, so it is a waste to store it in the cache. If you wait until it is requested a second time before you cache it, you avoid those single use pages going into your cache, and don’t hurt overall performance that much, because the content that is most useful to cache is requested a lot, and you only have to make one extra origin request.
- DeepSeaTortoise 7mo ago> Doing it on demand still utilizes their cached version, so it saves a trip to the origin, but doesn’t require doubling the cache size. They can still cache the results if the same site is scraped multiple times, but this saves having to cache things that are never going to be requested. Isn't this solving a slightly, but very significantly different problem? You could serve the very same data in two different ways: One to present to the users and one to hand over to scrapers. Of course, some sites would be too difficult or costly to transform into a common underlying cache format, but people who WANT their sides accessible to scrapers could easily help the process along a bit or serve their site in the necessary format in the first place. But the key is: A tool using a "pre-scraped" version of a site has very likely very different requirements of how a CDN caches this site. And this could be easily customizable by those using this endpoint. Want a free version? Ok, give us the list of all the sites you want, then come back in 10min and grab everything in one go, the data will be kept ready for 60s. Got an API token? 10 free near-real-time request for you and they'll recharge at a rate of 2 per hour. Want to play nice? Ask the CDN to have the requested content ready in 3 hours. Got deep pockets? Pay for just as many real-real-time requests as you need. What makes this so different is that unless customers are willing to hand over a lot of money, you dont need to cache anything to serve requests at all. Potentially not even later if you got enough capacity to serve the data for scheduled requests from the storage network directly. You just generate an immediate promise response to the request telling them to come back later. And depending on what you put into that promise, you've got quite a lot of control over the schedule yourself. - Got a "within 10min" request but your storage network has plenty if capacity in 30s? Just tell them to come back in 30s. - A customer is pushing new data into your network around 10am and many bots are interested in getting their hands on it as soon as possible, making requests for 10am to 10:05? Just bundle their requests. - Expected data still not around at 10:05? Unless the bots set an "immediate" flag (or whatever) indicating that they want whatever state the site is in right now, just reply with a second promise when they come back. And a third if necessary... and so on.
- hrmtst93837 7mo ago[flagged]
- Symbiote 7mo agoI think Common Crawl already offers this, although it's free: https://commoncrawl.org/ https://commoncrawl.org/
- brookst 7mo agoThat was my first thought when I read the headline. It would make perfect sense, and would allow some websites to have best of both worlds: broadcasting content without being crushed by bots. (Not all sites want to broadcast, but many do).
- Fokamul 7mo agoBut think about poor phishers and malware devs protected by Cloudflare.
- ryan14975 7mo agoThis makes a lot of sense. Cloudflare already has the rendered content at edge — serving a structured snapshot from cache would eliminate redundant crawling entirely. What I'd love to see is site owners being able to opt in and control the format. Something like a /cdn-cgi/structured endpoint that respects your robots.txt directives but gives crawlers clean markdown or JSON instead of making them parse raw HTML. The site owner wins (less bot traffic), the crawler wins (structured data), and Cloudflare wins (less load on origin).
- 8cvor6j844qw_d6 7mo agoDoes this bypass their own anti-AI crawl measures? I'll need to test it out, especially with the labyrinth.
- canpan 7mo agoI feel there is a conflict of interest here.. I'm split between: Yes! At last something to get CF protected sites! And: Uh! Now the internet is successfully centralized.
- xhcuvuvyc 7mo agoYeah, that'd be huge, like 90% of my search engine results are just cloudflare bot checks if I don't filter it out.
- mdasen 7mo agoIf this does bypass their own (and others') anti-AI crawl measures, it'd basically mean that the only people who can't crawl are those without money. We're creating an internet that is becoming self-reinforcing for those who already have power and harder for anyone else. As crawling becomes difficult and expensive, only those with previously collected datasets get to play. I certainly understand individual sites wanting to limit access, but it seems unlikely that they're limiting access to the big players - and maybe even helping them since others won't be able to compete as well.
- adi_kurian 7mo agoCommon Crawl has free egress
- jsheard 7mo agoThey say it doesn't: https://developers.cloudflare.com/browser-rendering/faq/#will-browser-rendering-be-detected-by-bot-management https://developers.cloudflare.com/browser-rendering/faq/#wil... Further down they also mention that the requests come from CFs ASN and are branded with identifying headers, so third party filters could easily block them too if they're so inclined. Seems reasonable enough.
- memothon 7mo agoI've used browser rendering at work and it's quite nice. Most solutions in the crawling space are kind of scummy and designed for side-stepping robots.txt and not being a good citizen. A crawl endpoint is a very necessary addition!
- Imustaskforhelp 7mo agoThis might be really great! I had the idea after buying https://mirror.forum https://mirror.forum recently (which I talked in discord and archiveteam irc servers) that I wanted to preserve/mirror forums (especially tech) related [Think TinyCoreLinux] since Archive.org is really really great but I would prefer some other efforts as well within this space. I didn't want to scrape/crawl it myself because I felt like it would feel like yet another scraping effort for AI and strain resources of developers. And even when you want to crawl, the issue is that you can't crawl cloudflare and sometimes for good measure. So in my understanding, can I use Cloudflare Crawl to essentially crawl the whole website of a forum and does this only work for forums which use cloudflare ? Also what is the pricing of this? Is it just a standard cloudflare worker so would I get free 100k requests and 1 Million per the few cents (IIRC) offer for crawling. Considering that Cloudflare is very scalable, It might even make sense more than buying a group of cheap VPS's Also another point but I was previously thinking that the best way was probably if maintainers of these forums could give me a backup archive of the forum in a periodic manner as my heart believes it to be most cleanest way and discussing it on Linux discord servers and archivers within that community and in general, I couldn't find anyone who maintains such tech forums who can subscribe to the idea of sharing the forum's public data as a quick backup for preservation purposes. So if anyone knows or maintains any forums myself. Feel free to message here in this thread about that too.
- ipaddr 7mo ago"I didn't want to scrape/crawl it myself because I felt like it would feel like yet another scraping effort for AI and strain resources of developers" You feel better paying someone to do the same thimg?
- Imustaskforhelp 7mo agoI actually don't but it seems that cloudflare caches responses so if anything instead of straining the developer resources, it would strain more cloudflare resources and cloudflare could better handle that more efficiently with their own crawl product. Also, I am genuinely open to feedback (Like a lot) so just let me know if you know of any other alternative too for the particular thing that I wish to create and I would love to have a discussion about that too! I genuinely wish that there can be other ways and part of the reason why I wrote that comment was wishing that someone who manages forums or knows people who do can comment back and we can have a discussion/something-meaningful! I am also happy with you also suggesting me any good use cases of the domain in general if there can be made anything useful with it. In fact, I am happy with transferring this domain to you if this is something which is useful to ya or anyone here (Just donate some money preferably 50-100$ to any great charity in date after this comment is made and mail me details and I am absolutely willing to transfer the domain, or if you work in any charity currently and if it could help the charity in any meaningful manner!) I had actually asked archive team if I could donate the domain to them if it would help archive.org in any meaningful way and they essentially politely declined. I just bought this domain because someone on HN said mirror.org when they wanted to show someone else mirror and saw the price of the .org domain being so high (150k$ or similar)and I have habit of finding random nice TLD and I found mirror.forum so I bought it And I was just thinking of hmm what can be a decent idea now that I have bought it and had thought of that. Obviously I have my flaws (many actually) but I genuinely don't wish any harm to anybody especially those people who are passionate about running independent forums in this centralized-web. I'd rather have this domain be expired if its activation meant harm to anybody. looking forward to discussion with ya.
- ljm 7mo agoIs cloudflare becoming a mob outfit? Because they are selling scraping countermeasures but are now selling scraping too. And they can pull it off because of their reach over the internet with the free DNS.
- rrr_oh_man 7mo ago[flagged]
- stri8ted 7mo agoDo you have any evidence to support this view?
- rolymath 7mo agoRead who and how it was founded. It's not a secret at all.
- rrr_oh_man 7mo agoIt’s funny how I got immediately downvoted and flagged
- pocksuppet 7mo agoWho else would MITM 30% of the internet?
- mtmail 7mo agoAny kind of source for the claim?
- Retr0id 7mo agoFor a long time cloudflare has proudly protected DDoS-as-a-service sites (but of course, they claim they don't "host" them)
- Dylan16807 7mo ago
- pupppet 7mo agoCloudflare getting all the cool toys. AWS, anyone awake over there?
- jppope 7mo agoThis is actually really amazing. Cloudflare is just skating to where the puck is going to be on this one.
- babelfish 7mo agoDidn't they just throw a (very public) fit over Perplexity doing the exact same thing?
- fleebee 7mo agoThe most egregious thing Perplexity did was to straight up ignore robots.txt. Cloudflare promise not to do that, so if we take their word for it, it's a quite different setup. That said, I'm not fan of letting users forge whatever user agents they please. Instead, AIUI to opt-out of getting crawled I have to look for the existence of certain request headers[1]. [1]: https://developers.cloudflare.com/browser-rendering/reference/automatic-request-headers/ https://developers.cloudflare.com/browser-rendering/referenc...
- everfrustrated 7mo agoWill this crawler be run behind or infront of their bot blocker logic?
- shadowfiend 7mo agoIn front: https://developers.cloudflare.com/browser-rendering/rest-api/crawl-endpoint/#robotstxt-and-bot-protection https://developers.cloudflare.com/browser-rendering/rest-api...
- deleted 7mo ago[deleted]
- greatgib 7mo agoAll what was expected, first they do a huge campaign to out evil scrapers. We should use their service to ensure your website block LLMs and bots to come scraping them. Look how bad it is. And once that is well setup, and they have their walled garden, then they can present their own API to scrape websites. All well done to be used by your LLM. But as you know, they are the gate keeper so that the Mafia boss decide what will be the "intermediary" fee that is proper for itself to let you do what you were doing without intermediary before.
- shadowfiend 7mo agoNo: https://developers.cloudflare.com/browser-rendering/rest-api/crawl-endpoint/#robotstxt-and-bot-protection https://developers.cloudflare.com/browser-rendering/rest-api...
- x0x0 7mo agomost websites, particularly those behind cloudflare, are very restrictive even to crawlers that obey robots. Proof: a ton of my time over the last year, and my crawlers very carefully obey robots. It's hard to see how this isn't extorting folks by offering a working solution that, oh, cloudflare doesn't block. As long as you pay Cloudflare. Perhaps I'm overly cynical, but I'd be quite surprised if cloudflare subjected their own headless browsing to the same rules the rest of the internet gets.
- gruez 7mo ago>most websites, particularly those behind cloudflare, are very restrictive even to crawlers that obey robots. Proof: a ton of my time over the last year, and my crawlers very carefully obey robots. The docs are pretty equivocal though: >If you use Cloudflare products that control or restrict bot traffic such as Bot Management, Web Application Firewall (WAF), or Turnstile, the same rules will apply to the Browser Rendering crawler. It's not just robots.txt. Most (all?) restrictions that apply to outside bots apply to cloudflare's bot as well, at least that's what they're claiming. If they're being this explicit about it, I'm willing to give them the benefit of the doubt until there's evidence to the contrary, rather than being a cynic and assuming the worst.
- binarymax 7mo agoReally hard to understand costs here. What is a reasonable pages per second? Should I assume with politeness that I'm basically at 1 page per second == 3600 pages/hour? Seems painfully slow.
- devnotes77 7mo ago[flagged]
- devnotes77 7mo ago[flagged]
- zyz 7mo ago> Browser Rendering is only available on the Workers Paid plan ($5/month). It is not part of the free tier. The post says it's available for both free and paid plans. According to the pricing page of the Browser Rendering, the free plan will have 10 minutes/day browsing time.
- gingerlime 7mo ago[0] seems to suggest even paid plans are effectively limited to 500 web pages per day, right? Crawl jobs per day 5 per day Maximum pages per crawl 100 pages [0] https://developers.cloudflare.com/browser-rendering/limits/#crawl-endpoint-limits https://developers.cloudflare.com/browser-rendering/limits/#...
- patchnull 7mo ago[flagged]
- radium3d 7mo agoInstead of "should have been an email" this is "should have been a prompt" and can be run locally instead. There are a number of ways to do this from a linux terminal. ``` write a custom crawler that will crawl every page on a site (internal links to the original domain only, scroll down to mimic a human, and save the output as a WebP screenshot, HTML, Markdown, and structured JSON. Make it designed to run locally in a terminal on a linux machine using headless Google Chrome and take advantage of multiple cores to run multiple pages simultaneously while keeping in mind that it might have to throttle if the server gets hit too fast from the same IP. ``` Might use available open source software such as python, playwright, beautifulsoup4, pillow, aiofiles, trafilatura
- Normal_gaussian 7mo agoThis presumably is going to be cheap and effective. Its much easier to wrap a prompt round this and know it works that mess around with crawling it all yourself. You'll still be hand-rolling it if you want to disrespect crawling requirements though.
- supermdguy 7mo agoI’ve actually written a crawler like that before, and still ended up going with Firecrawl for a more recent project. There’s just so many headaches at scale: OOMs from heavy pages, proxies for sites that block cloud IPs, handling nested iframes, etc.
- Keyframe 7mo agoThat'd be more like that draw an owl meme. Devil's in the details. Holy shit, there's so many details...
- skybrian 7mo agoIf two customers crawl the same website and it uses crawl-delay, how does it handle that? Are they independent, or does each one run half as fast?
- PeterStuer 7mo agoYou put a governor on the domain, and you return from the cache instead.
- arjie 7mo agoOh man, I was hoping I could offer a nicely-crawled version of my site. It would be cool if they offered that for site admins. Then everyone who wanted to crawl would just get a thing they could get for pure transfer cost. I suppose I could build one by submitting a crawl job against myself and then offering a `static.` subdomain on each thing that people could access. Then it's pure HTML instant-load.
- echoangle 7mo agoI don’t really get the usecase. Is your site static? Then you should just render it to html files and host the static files. And if it’s not static, how would a snapshot of the pages help if they change later? And also why not just add some caching to the site then?
- arjie 7mo agoAh the use-case is archive.org but fast. But it's okay. Before I die I will make the static copy of my site myself.
- david_iqlabs 7mo ago[flagged]
- Normal_gaussian 7mo ago"Well-behaved bot - Honors robots.txt directives, including crawl-delay" From the behaviour of our peers, this seems to be the real headline news.
- arjunchint 7mo agoRIP @FireCrawl or at the very least they were the inspiration for this?
- ed_mercer 7mo ago> Honors robots.txt directives, including crawl-delay Sounds pretty useless for any serious AI company
- PeterStuer 7mo agoWhat % of sites have a content update volume that exceeds what you can get respecting crawl delay? If your delay is 1s and you publish less than 60 updates a minute on average I can still get 100%. Most crawls are not that latency sensitive, certainly not the ai ones. HFT bots, now that is an entirely different ballgame.
- mrweasel 7mo ago> Most crawls are not that latency sensitive, certainly not the ai ones. They certainly behave like they are. We constantly see crawlers trying to do cache busting, for pages that hasn't change in days, if not weeks. It's hard to tell where the bots are coming from theses days, as most have taken to just lie and say that they are Chrome. I'd agree that the respecting robots.txt makes this a non-starter for the problematic scrapers. These are bots that that will hammer a site into the ground, they don't respect robots.txt, especially if it tells them to go away. All of this would be much less of a problem if the authors of the scrapers actually knew how to code, understood how the Internet works and had just the slightest bit of respect for others, but they don't so now all scrapers are labeled as hostile, meaning that only the very largest companies, like Google, get special access.
- moebrowne 7mo ago> We constantly see crawlers trying to do cache busting Do you have a source for this? Not saying you're wrong, I'd just like to know more
- mrweasel 7mo agoNot really, given that the work we do in that direction isn't exactly public. You can recreate the scenario though. Spin up a wiki of some sort, scrapers love wikis, ideally enable some form of caching, and just sit back and watch scrapers throw random shit in the URL parameters.
- Lasang 7mo agoThe idea of exposing a structured crawl endpoint feels like a natural evolution of robots.txt and sitemaps. If more sites provided explicit machine-readable entry points for crawlers, indexing could become a lot less wasteful. Right now crawlers spend a lot of effort rediscovering the same structure over and over. It also raises interesting questions about whether sites will eventually provide different views for humans vs. automated agents in a more formalized way.
- catlifeonmars 7mo ago> It also raises interesting questions about whether sites will eventually provide different views for humans vs. automated agents in a more formalized way. This question raises an interesting question about if this would exacerbate supply chain injection attacks. Show the innocuous page to the human, another to the bot.
- _heimdall 7mo agoI expect that if we still used REST indexing would be even less wasteful. I've found myself falling pretty hard on the side of making APIs work for humans and expecting LLM providers to optimize around that. I don't need an MCP for a CLI tool, for example, I just need a good man page or `--help` documentation.
- rglover 7mo agoI just do a query param to toggle to markdown/text if ?llm=true on a route. Easy pattern that's opt-in.
- pdntspa 7mo agoThey already do... A lot of known crawlers will get a crawler-optimized version of the page
- rafram 7mo agoDo they? AFAIK Google forbids that, and they’ll occasionally test that you aren’t doing it.
- tjpnz 7mo agoDo I have the option to fill it with junk for LLMs?
- pqdbr 7mo agoOff-topic, but I'm having a terrible experience with Cloudflare and would love to know if someone could offer some help. All of a sudden, about 1/3 of all traffic to our website is being routed via EWR (New York) - me included -, even tough all our users and our origin servers are in Brazil. We pay for the Pro plan but support has been of no help: after 20 days of 'debugging' and asking for MTRs and traceroutes, they told us to contact Claro (which is the same as telling me to contact Verizon) because 'it's their fault'.
- tgrowazay 7mo agoIt is possible that Claro has a bad route that sends all traffic destined for Cloudflare through New York.
- tempest_ 7mo agoEvery once and a while we have had Bell Canada route a request that should be going about 6 blocks away across the continent and back. They are not super helpful fixing it either.
- weird-eye-issue 7mo agoDo you think cloudflare is responsible for all of the network traffic routing in the entire world and can simply fix any problem even if it's on somebody else's network?
- pqdbr 7mo agoNo. I do think that Cloudflare is a great company and got where it's at today because they care for this type of issue, and has a much better chance of contacting their peering traffic partner than me because they take care of ~20% of all internet traffic, while I take care of none.
- fbrncci 7mo agoAwesome, so I no longer have to use Firecrawl or my own crawler to scrape entire websites for an agent? Especially when needing residential proxies to do so on Cloudflare protected sites? Why though?
- freakynit 7mo agoI have tried theirs... they are NOT proxies.. that means majority of the popular sites actually block scraping... even if they are protected by cloudflare itself.
- coreq 7mo agoThe big question here is this a verified-bot on the Cloudflare WAF? Didn't Google get into trouble for using their search engine user agent and IPs to feed Gemini in Europe?
- charcircuit 7mo ago>Honors robots.txt Is it possible to ignore robot.txt in the case the crawl was triggered by a human?
- 1vuio0pswjnm7 7mo agoCan a CDN be a "walled garden"
- sourcecodeplz 7mo agoI love this from CloudFlare!
- ramblurr 7mo agoIt seems like there's a missed use case: web archiving. I don't see any mention of WARC as an output format. This could be useful to journalists and academically if they had it.
- mrexcess 7mo agoGETOLD /index.html "2026-03-11T10:30:45Z" would be such cool functionality...
- rixrax 7mo agoAnd while at it, ability to mount the resulting archives at some virtual root in nginx|apache. E.g. serve site-archive.extension at /somepath/site. And standalone simple webserver that can take one or more archives from command line and servers them.
- RamblingCTO 7mo agoDoesn't work for pages protected by cloudflare in my experience. What a shame, they could've produced the problem and sold the solution.
- chvid 7mo agoAs long at it gets past Azure's bot protection ...
- GodelNumbering 7mo agoI imagine that would cause a backlash from the website owners trusting cloudflare to keep their content 'safe'
- davidhariri 7mo agoCame here to write this. I am getting much better results from Firecrawl (not affiliated with them, just a happy customer).
- RamblingCTO 7mo agofuck firecrawl. they copied my idea by showing interest in my product and then copied it, used their YC money to give it all out for free. fuck nick in particular. I'm still salty over this
- neversupervised 7mo agoTell more. Crawling is not a new idea. How did they abuse you?
- xeornet 7mo ago"they copied my idea by showing interest in my product and then copied it". What exactly is revolutionary about Firecrawl or your product? Scraping APIs have been around for over a decade.
- RamblingCTO 7mo ago
- allixsenos 7mo ago"Selling the wall and the ladder." "Biggest betrayal in tech." "Protection racket." These hot takes sound smart but they're not. The web was built to be open and available to everyone. Serving static HTML from disk back in the day, nobody could hurt you because there was nothing to hurt. We need bot protection now because everything is dynamic, straight from the database with some light caching for hot content. When Facebook decides to recrawl your one million pages in the same instant, you're very much up shit creek without a paddle. A bot that crawls the full site doesn't steal anything, but it does take down the origin server. My clients never call me upset that a bot read their blog posts. They call because the bot knocked the site offline for paying customers. Bot protection protects availability, not secrecy. And the real bot problem isn't even crawling. It's automated signups. Fake accounts messaging your users. Bots buying out limited drops before a human can load the page. Like-farming. Credential stuffing. That's what bot protection is actually for: preventing fraud, not preventing someone from reading your public website. Cloudflare's `/crawl` respects robots.txt. Don't want your content crawled, opt out. But if you want it indexed and can't handle the traffic spike, this gets your content out without hammering production. As for the folks saying Cloudflare should keep blocking all crawlers forever: AI agents already drive real browsers. They click, scroll, render JavaScript. Go look at what browser automation frameworks can do today and then explain to me how you tell a bot from a person. That distinction is already gone. The hot takes are about a version of the internet that doesn't exist anymore.
- iranu 7mo agoHonestly, it feels like cloudflare bullying other sites into using their anti-bot services. great business model by charging owners and devs at the same time. Using AI per page to parse content. its reckless.
- radicalriddler 7mo agoInteresting... I built an MCP server for their initial browser render as markdown, and I just tell the LLM to follow reasonable links to relative content, and recurse the tool.
- ClaudeAgent_WK 7mo ago[flagged]
- branoco 7mo agoInterested how this unfolds
- kelvinjps10 7mo agoCould they collaborate with the website's creators that have websites behind cloudfare to allow their content to be accessed via an API in exchange of a compensation?. This could be a way to compensate creators and AI companies be able to access content that's unreachable as it's protected by cloudfare
- lathiat 7mo agoThey are one step ahead of you: https://blog.cloudflare.com/introducing-pay-per-crawl/ https://blog.cloudflare.com/introducing-pay-per-crawl/ Sort of though. Still private beta since July 2025.
- carloslfu 7mo agoThey have a Pay Per Crawl option for owners. This plus a /crawl endpoint is genius.
- xorgun 7mo ago[dead]
- superkuh 7mo agoCloudflare are mafiosos. They create the problem and then sell you the solution to themselves.
- bobpaw 7mo agoTIL about the Crawl-delay directive. Although it seems that most honest bots move slower and dishonest bots will learn to.
- andrethegiant 7mo agoI tried to make exactly this a year ago. Built on Cloudflare using all of their primitives: https://crawlspace.dev https://crawlspace.dev -- It didn't work too well (so don't bother trying it).
- keeda 7mo agoIf anyone is taking feature requests, could you add an option to return the snapshot as an MHTML with all static assets embedded? (I know this could get inefficient from a storage perspective, but if it really matters you could dedupe assets on your end, which is what my janky homegrown crawler does.)
- stevenhubertron 7mo agoQueue-It protected pages catch it as well and prevent crawling.
- kordlessagain 7mo agoFuck Cloudflare.
- kseniamorph 7mo agoI remember reading a CF blog post about crawler separation and responsible AI bot principles where they argue every bot should have one distinct purpose. Now they're building crawling infrastructure themselves, and their own /crawl endpoint lists "training AI systems" as a use case alongside regular crawling. So not only are they in the crawling business now, they're not following the separation principle. To be fair, there's a business logic here. But it's hard not to notice the irony. https://blog.cloudflare.com/uk-google-ai-crawler-policy/ https://blog.cloudflare.com/uk-google-ai-crawler-policy/
- sireat 7mo agoOne has to be highly suspicious of any "fair, better for others" claims coming from corporate entities. It is the ages old story of https://en.wikipedia.org/wiki/Quod_licet_Iovi%2C_non_licet_bovi https://en.wikipedia.org/wiki/Quod_licet_Iovi%2C_non_licet_b... Also brings back the irony now apparent in original Google paper: http://infolab.stanford.edu/pub/papers/google.pdf http://infolab.stanford.edu/pub/papers/google.pdf "To make matters worse, some advertisers attempt to gain people’s attention by taking measures meant to mislead automated search engines."
- laalshaitaan 7mo agoIMO the under-discussed risk here is that sites will start serving different content to verified crawlers vs real users. You're already seeing it with known search bots getting sanitized views. If your agent's context comes from a crawl the site knows is going to an AI, you have no guarantee it matches what a human sees, and that data quality problem won't surface until your agent starts acting on selectively curated information. This could go wrong on same levels.
- vimda 7mo agoThis already happens in the opposite direction. See: news websites that drop their pay wall for GoogleBot
- devnotes77 7mo ago[flagged]
- m3047 7mo agoSeems like it was just hours ago they started reaching out to my edge servers from their address space (Me: why is a reverse proxy service banging my servers when I'm not a customer? did some miscreant sign me up somehow?) and it was for Apple, privacy, mom and pie (a VPN service, dressed in noble aspirations). It never quite smelled like pie to me. If you're doing threat hunting / risk enumeration, Cloudflare is no longer a passive service that miscreants hide behind, they now actively reach out and grab your privates.
- ryan14975 7mo ago[flagged]
- nathanhouse 7mo agoI built a CLI wrapper for the Browser Rendering REST API — covers all 9 endpoints including /crawl. Two Bun scripts, zero dependencies: one for single-page ops (render, screenshot, PDF, scrape, AI extraction), one for multi-page crawling. Also works as Claude Code slash commands if you're into that. https://github.com/nathanhouse/cloudflare-browser-rendering-cli https://github.com/nathanhouse/cloudflare-browser-rendering-...
- Ian_Kerins 7mo agoA lot of the discussion around the /crawl endpoint seems to miss a key detail in the docs. The crawler explicitly identifies itself as a bot, respects robots.txt, and does not bypass CAPTCHAs, WAF rules, or Cloudflare Bot Management. So technically it’s a nice managed crawling system, but in practice it only works on sites that already allow bots to crawl them. For many real-world data extraction use cases, the problem isn’t crawling infrastructure, it’s dealing with sites that actively block bots. In those cases you still need traditional scraping approaches.
- iwinux 6mo agoCloudflare: pay me to keep crawlers away Also Cloudflare: pay me to get your crawlers through my anti-crawler firewall