4 ms·
Details here: https://blog.cloudflare.com/introducing-pay-per-crawl/ https://blog.cloudflare.com/introducing-pay-per-crawl/
by yablak 1y ago
Details here: https://blog.cloudflare.com/introducing-pay-per-crawl/ https://blog.cloudflare.com/introducing-pay-per-crawl/
- grg0 1y agoThis requires good faith on behalf of the crawler? So it's DOA; why even bother implementing this? Also, what a piece of zero-trust shit the web is becoming thanks to a couple of shit heads who really need to extract monetary value out of everything. Even if this non-solution were to work, the prospect of putting every website behind Cloudsnare is not a good one anyway. What the web needs right now, to be honest, is machetes. In ample quantity. Tell me who's running that crawler that is bothering you and I will put them to the sword. They won't even need to present a JWK in the header.
- xg15 1y agoMaybe I didn't understand the proposal completely yet, but wouldn't the crawler only have to cooperate (send the right headers, implement that auth framework, etc) if they want to pay? The standard response to a crawler is a 402 Payment Required response, probably as a result of an aggressive bot detection. So essentially, it's turning a site's entire content into an API: Either sign up for an API key or get blocked. The question remains though how well they will be able to distinguish bot traffic from humans - also, will they make an exception for search engines?
- grg0 1y agoThat is not what I understood, and it sounds terrible. What if you're not a crawler but random Joe surfing the internet? Clearly Joe should see content without payment? So they need some way to tell the crawler and Joe apart, and presumably they require the crawler to set certain request headers. The headers aren't just to issue the payment, it's to identify the crawler in the first place?
- xg15 1y agoThe writeup doesn't talk about actively misbehaving crawlers a lot, but this bit implies for me that the headers are for the "happy path", I.e. crawlers wanting to pay: > Each time an AI crawler requests content, they either present payment intent via request headers for successful access (HTTP response code 200), or receive a 402 Payment Required response with pricing. I don't see how it would make sense otherwise, as the requirements for crawlers include applying for a registration with Cloudflare. Who in their right mind would jump through registration hoops only so they can not access a site? This wouldn't even keep away the crawlers that are operating today. I agree there has to be some way to distinguish crawlers from regular users, but the only way I can see how this could be done is with bot detection algorithms. ...which are imperfect and will likely flag some legitimate human users as bots. So yes, this will probably leading to web browsing becoming even more unpleasant.
- quixoticaxolotl 1y agoJoe will be fine. Cloudflare is pretty good at differentiating humans from bot traffic - see how we do it here: https://developers.cloudflare.com/turnstile/ https://developers.cloudflare.com/turnstile/ The idea behind the headers is to allow bots to bypass automatic bot filtering, not blockade all regular traffic. In other words: - we block bots (the website owner can configure how aggressively we block) - unless they say they're from an AI crawler we've vetted, as attested by the signature headers - in which case we let them pay - and then they get to access the content (Disclosure: I wrote the web bot auth implementation Cloudflare uses for pay per crawl)
- cryptonector 1y agoIt's Cloudflare. That means they are good at DoS and DDoS protection. AI crawlers are basically DoS agents. I think CF can start with an honor system that also has attached to it the implied threat to block crawlers from all CF hosted content, and that is a pretty big hammer to hit the abusers with. So I'm cautiously optimistic. Well, I suppose pessimistic too: if this works what this will mean is that all contents will end up moving into big player hosting like CF.