28 ms·
The web does not need gatekeepers: Cloudflare’s new “signed agents” pitch
- Havoc 1y agoDon't think the base game plan here is necessarily all that bad. It being concentrated in one for profit entity however very much is
- cyberlurker 1y agoI would love that vision to become reality but what Cloudflare is doing is unfortunately necessary atm.
- fooqux 1y agoOk, I'll bite. Why is turning the Internet into a walled garden necessary now?
- giantrobot 1y agoMulti-Tbps DDoS attacks, pervasive scanning of sites for exploits, comically expensive egress bandwidth on services like AWS, and ISPs disallowing hosting services on residential accounts.
- doublerabbit 1y agoStart forcing tighter security on the devices causing the Multi-Tbps DDoS attacks would be a better option, no? Cheap unsecured IoT devices are a problem. It's not just computers anymore. Web enabled CCTV, doorbell cameras are all culprits.
- jurking_hoff 1y ago[dead]
- esseph 1y agoCommercial, criminal, and state interests have far more resources than you do, and their interests are in direct conflict with yours. That would be fine, you could walk away and go home, but if you're going to drive on their digital highways, you're going to need "insurance" just protect you from everyone else. Ongoing multi-nation WWIII-scale hacking and infiltration campaigns of infrastructure, AI bot crawling, search company and startup crawling, security researchers crawling, and maybe somebody doesn't like your blog and decides to rent a botnet for a week or so. Bet your ISP shuts you off before then to protect themselves. (Happens all the time via BGP blackholing, DDoS scrubbing services, BGP FlowSpec, etc).
- jjangkke 1y agoI would love to get off Cloudflare but there are no real good alternatives
- didibus 1y agoAWS is an alternative no?
- miohtama 1y agoAWS needs a dedicated AWS engineer while any technical person and some non-technical people have skill to set up Cloudflare. Esp. Without surprise bills.
- didibus 1y agoI always hear this, but honestly I'm not sure it's true. It's hard to assess the validity of this versus Cloudflare having a really good marketing department. I've used neither, so I can't say, but I've also never seen anyone truly explain why/why-not.
- ryoshu 1y agoWhy not use both and find out? Cloudflare is much less technical than AWS, but still a bit technical.
- hdgvhicv 1y agoI thought the whole point of paying a fortune for AWS was to avoid having a dedicated engineer. It’s the cobol of the 21st century.
- nromiun 1y agoBankruptcy as a surprise gift is not an alternative. Even those that use big cloud providers like AWS and GCP use CDNs like Cloudflare to protect themselves. And there is no free CDN like Cloudflare.
- matt-p 1y agoI have zero issue with Ai Agents, if there's a real user behind there somewhere. I DO have a major issue with my sites being crawled extremely aggressively by offenders including Meta, Perplexity and OpenAI - it's really annoying realising that we're tying up several cpu cores on AI crawling. Less than on real users and google et al.
- Operyl 1y agoThey're getting to the point of 200-300RPS for some of my smaller marketing sites, hallucinating URLs like crazy. It's fucking insane.
- matt-p 1y agoI'm seeing around the same, as a fairly constant base load. Even more annoying when it's hitting auth middleware constantly, over and over again somehow expecting a different answer.
- palmfacehn 1y agoYou'd think they would have an interest in developing reasonable crawling infrastructure, like Google, Bing or Yandex. Instead they go all in on hosts with no metering. All of the search majors reduce their crawl rate as request times increase. On one hand these companies announce themselves as sophisticated, futuristic and highly-valued, on the other hand we see rampant incompetence, to the point that webmasters everywhere are debating the best course of action.
- matt-p 1y agoHonestly it's just tragedy of the commons. Why put the effort in when you don't have to identify yourself, just crawl and if you get blocked move the job to another server.
- palmfacehn 1y agoAt this point I'm blocking several ASNs. Most are cloud provider related, but there are also some repurposed consumer ASNs coming out of the PRC. Long term, this devalues the offerings of those cloud providers, as prospective customers will not be able to use them for crawling.
- sdsd 1y agoMaybe the title means something more like "The web should not have gatekeepers (Cloudflare)". They do seem to say as much toward the end: >We need protocols, not gatekeepers. But until we have working protocols, many webmasters literally do need a gatekeeper if they want to realistically keep their site safe and online. I wish this weren't the case, but I believe the "protocol" era of the web was basically ended when proprietary web 2.0 platforms emerged that explicitly locked users in with non-open protocols. Facebook doesn't want you to use Messenger in an open client next to AIM, MSN, and IRC. And the bad guys won. But like I said, I hope I'm wrong.
- jeroenhd 1y ago>We need protocols, not gatekeepers The funny thing is that this blog post is complaining about a proposed protocol from Cloudflare (one which will identify bots so that good bots can be permitted). The signup form is just a method to ask Cloudflare (or any other website owner/CDN) to be categorized as a good bot. It's not a great protocol if you're in the business of scraping websites or selling people bots to access websites for them, but it's a great protocol for people who just want their website to work without being overwhelmed by the bad side of the internet. The whitelist approach Cloudflare takes isn't good for the internet, but for website owners who are already behind Cloudflare, it's better than the alternative. Someone will need to come up with a better protocol that also serves the website owners' needs if they want Cloudflare to fail here. The AI industry simply doesn't want to cooperate, so their hand must be forced, and only companies like Cloudflare are powerful enough to accomplish that.
- ccgreg 1y agoConventional crawlers already have a way to identify themselves, via a json file containing a list of IP addresses. Cloudflare is fully aware of this defacto standard.
- jimmyl02 1y agoI understand the concerns around a central gatekeeper but I'm confused as to why this specifically is viewed negatively. Don't website owners have to choose to enable cloudflare and to opt-in to this gate that the site owners control? If this was cloudflare going into some centralized routing of the internet and saying everything must do X then that would be a lot more alarming but at the end of the day the internet is decentralized and site owners are the ones who are using this capability. Additionally I don't think that I as an individual website owner would actually want / be capable of knowing which agents are good and bad and cloudflare doing this would be helpful to me as a site owner as long as they act in good faith. And the moment they stop acting in good faith I would be able to disable them. This is definitely a problem right now as unrestricted access to the bots means bad bots are taking up many cycles raising costs and taking away resources from real users
- jmarbach 1y agoI recently ran a test on the page load reliability of Browserbase and I was shocked to see how unreliable it was for a standard set of websites - the top 100 websites in the US by traffic according to SimilarWeb. 29% of page load requests failed. Without an open standard for agent identification, it will always be a cat and mouse game to trap agents, and many agents will predictably fail simple tasks. https://anchorbrowser.io/blog/page-load-reliability-on-the-top-100-websites-in-the-us https://anchorbrowser.io/blog/page-load-reliability-on-the-t... Here's to working together to develop a new protocol that works for agents and website owners alike.
- jmtame 1y agoI pretty much use Perplexity exclusively at this point, instead of Google. I'd rather just get my questions answered than navigate all of the ads and slowness that Google provides. I'm fine with paying a small monthly fee, but I don't want Cloudflare being the gatekeeper. Perhaps a way to serve ads through the agents would be good enough. I'd prefer that to be some open protocol than controlled by a company.
- verdverm 1y agoPerplexity has been one of the AI companies that created the problem that gave rise to this CF proposal. Why doesn't Perplexity invest more into being a responsible scraper? https://blog.cloudflare.com/perplexity-is-using-stealth-undeclared-crawlers-to-evade-website-no-crawl-directives/ https://blog.cloudflare.com/perplexity-is-using-stealth-unde...
- jmtame 1y agoRe-read what I wrote.
- verdverm 1y agoand what am I supposed to garner from the re-read? What did you say that relates to Perplexity being one of the reasons that Cloudflare and their customers have decided they need better protection from abusive scrapers? Websites choose their own gatekeepers, Cloudflare is just one provider
- Fabricio20 1y agoThis has been my experience more recently as well, I've finally migrated from google to Brave Search since google was just slow for me. I also appreciate the AI search results a bit when im looking for something very specific (like what the yaml definition for a docker swarm deployment constraint looks like) because the AI just gives me the snippet while the search results are 300 medium blog posts about how to use docker and none of them explain the variables/what each does. Even the official docker documentation website is a mess to navigate and find anything relevant!
- johnnienaked 1y agoI can see a future where I don't use the internet at all.
- robmusial 1y agoMaybe not the Internet for me, but certainly the web. But I totally agree with the sentiment.
- theideaofcoffee 1y agoYour ideas are intriguing to me and wish to subscribe to your newsletter. Joking aside, I think the ideas and substance are great and sorely needed. However, I can only see the idea of a sort of token chain verification as running into the same UX problems that plagued (plagues?) PGP and more encryption-focused processes. The workflow is too opaque, requires too much specialized knowledge that is out of reach for most people. It would have to be wrapped up into something stupid simple like an iOS FaceID modal to have any hope of succeeding with the general public. I think that's the idea, that these agents would be working on behalf of their owners on their own devices, so it has to be absolutely seamless. Otherwise, rock on.
- impure 1y agoWell, if you have a better way to solve this that’s open I’m all ears. But what Cloudflare is doing is solving the real problem of AI bots. We’ve tried to solve this problem with IP blocking and user agents, but they do not work. And this is actually how other similar problems have been solved. Certificate authorities aren’t open and yet they work just fine. Attestation providers are also not open and they work just fine.
- viktorcode 1y agoAI poisoning is a better protection. Cloudflare is capable of serving stashes of bad data to AI bots as protective barrier to their clients.
- esseph 1y agoAI poisoning is going to get a lot of people killed, be cause the AI won't stop being used.
- viktorcode 1y agoBy that logic AI already killing people. We can't presume that whatever can be found on the internet is reliable data, can't we?
- lucb1e 1y agoIf science taught us anything it's that no data is ever reliable. We are pretty sure about so many things, and it's the best available info so we might as well use it, but in terms of "the internet can be wrong" -> any source can be wrong! And I'd not even be surprised if internet in aggregate (with the bot reading all of it) is right more often than individual authors of pretty much anything
- esseph 1y agoYet we use it every day for police, military, and political targeting with economic and kinetic consequences.
- skybrian 1y agoThis is sort of like how email is based on Internet standards but a large percentage of email users use Gmail. The Internet standards Cloudflare is promoting are open, but Cloudflare has a lot of power due to having so many customers. (What are some good alternatives to Cloudflare?) Another way the situation is similar: email delivery is often unreliable and hard to implement due to spam filters. A similar thing seems to be happening to the web.
- nromiun 1y agoIt is a big problem. There is no good alternative to Cloudflare as a free CDN. They put servers all over the world and they are giving them away for free. And making their money on premium serverless services. Not to mention the big cloud providers are unhinged with their egress pricing.
- gck1 1y ago> Not to mention the big cloud providers are unhinged with their egress pricing. I always wonder why this status quo persisted even after Cloudflare. Their pricing is indeed so unhinged, that they're not even in consideration for me for things where egress is a variable. Why is egress seemingly free for Cloudflare or Hetzner but feels like they launch spaceships at AWS and GCP every time you send a data packet to the outside world?
- nromiun 1y agoThey are just greedy. And they know nobody can compete with them for availability in every country. Except for Cloudflare, which is why it is so popular.
- whimsicalism 1y ago[flagged]
- positiveblue 1y agoNarrator: but he did put effort... Anyway, main take aways for you: - We REALLY need a way to tie identity to agent/requests - The idea of registering with cloudflare to be able to access a website is bad - Sites should be able to block whoever they want or make anything they desire a requirement (there are people that has login only with google after all) We have the right primitives to build something that works for any provider (from cloudflare, to Akamai, to self hosting nginx servers). Let's take that route
- esseph 1y agoTying identity of any thing to this is a road we should not go down. I hear you, but this will get extended to people, which is already under threat in a lot of places.
- immibis 1y agoWe have a way to tie identity to agent/requests. It's called IP address. Every IP address has a provider's identity, and most providers can further specify an individual or business. It doesn't help. Do you know why it doesn't help? It doesn't help because AI companies just pay people to borrow your identity. Somewhere between $2 and $250 per month to borrow your internet identity - the former for running a program in the background on your computer and the latter for getting an internet contract in your name. Or you can pay developers of free mobile games to put your code in their game, and the user doesn't know about it, and it's not even illegal. All that will happen if you tighten identity further is that these measures will go further, e.g. farms of real devices with robots giving touch inputs and cameras running OCR. Meanwhile you'll make it more annoying for everyone to use the internet until they quit. You know what other effect it has? Only big players will be able to afford the workaround measures. Google will have no problem deploying an army of robots with phones (or just faking it since they own the signing keys) meanwhile you just banned every non-Google non-Apple phone from your website, forcing centralisation and monopoly on everyone. And for what? Saving 3 requests per second of server load?
- TIPSIO 1y agoEveryone loves the dream of a free for all and open web. But the reality is how can someone small protect their blog or content from AI training bots? E.g.: They just blindly trust someone is sending Agent vs Training bots and super duper respecting robots.txt? Get real... Or, fine what if they do respect robots.txt, but they buy the data that may or may not have been shielded through liability layers via "licensed data"? Unless you're reddit, X, Google, or Meta with scary unlimited budget legal teams, you have no power. Great video: https://www.youtube.com/shorts/M0QyOp7zqcY https://www.youtube.com/shorts/M0QyOp7zqcY
- anigbrowl 1y agoMaybe this is a naive question, but why not just cut an IP off temporarily if it sends too many requests or sends them too fast?
- ceejayoz 1y agoThey use many IPs, often not identifiable as the same bot.
- clvx 1y agoYou can lock it up with a user account and payment system. The fact the site is up on the internet doesn’t mean you can or cannot profit from it. It’s up to you. What I would like it’s a way to notify my isp and say, block this traffic to my site.
- inetknght 1y ago> What I would like it’s a way to notify my isp and say, block this traffic to my site. I would love that, and make it automated. A single message from your IP to your router: block this traffic. That router sends it upstream, and it also blocks it. Repeat ad nauseum until source changes ASN or (if the originator is on the same ASN) reaches the router from the originator, routing table space notwithstanding. Maybe it expires after some auto-expiry -- a day or month or however long your IP lease exists. Plus, of course, a way to query what blocks I've requested and a way to unblock.
- sugarpimpdorsey 1y agoThis is like saying companies don't need security gates and checkpoints. Unfortunately the world is filled with bad people, and you need security to keep them off your property.
- bendigedig 1y agoIf the broader economic system wasn't based on what is essentially theft, security wouldn't be as necessary as it is.
- ACCount37 1y agoWe have far too many gatekeepers as it is. Any attempt to add any more should be treated as an act of aggression. Cloudflare seems very vocal about its desire to become yet another digital gatekeeper as of late, and so is Google. I want both reduced to rubble if they persist in it.
- jeroenhd 1y agoSeveral companies are looking to provide a solution for the AI bot problem. Cloudflare stands to make a lot of money if people pick their solution. But Cloudflare backing down won't make the problem go away, and someone else's bad solution will be chosen instead. The gatekeeping described here is gatekeeping a website owner chooses. It's an alternative to pay walls, bespoke bot detection, or some kind of ID verification. Cloudflare already provides a service, but standardising the service will open up the market (at the cost of competitors adopting Cloudflare's standard). The freedom of the open web also extends to the owners of the websites people visit.
- account42 1y agoYou could use the same argument against antitrust laws. All monopolies were "chosen" by their customers, that doesn't mean we should give them a free pass.
- encom 1y agoWhat do you mean Google "desires" to become a gatekeeper? They have been a gatekeeper for years, since they control the browser everyone uses, and Firefox usage is now in the noise. Google just steers the www where they want it to go. Killing ublock, pushing .webp trash, etc.
- timshell 1y agoI think about this as a startup founder building a 'proof-of-human' layer on the Internet. One of the hard parts in this space is what level of transparency should you have. We're advancing the thesis that behavioral biometrics offers robust continuous authentication that helps with bot/human and good/bad, but people are obviously skeptical to trust black-box models for accuracy and/or privacy reasons. We've defaulted to a lot of transparency in terms of publishing research online (and hopefully in scientific journals), but we've seen the downside: competitors fake claims about their own best in-house behavioral tools that is behind their company walls in addition to investors constantly worried about an arms race. As someone genuinely interested (and incentivized!) to build a great solution in this space, what are good protocols/examples to follow?
- avtar 1y ago> An allowlist run by ONE company? An allowlist run by one company that site owners chose to engage with. But the irony of taking an ideological stance about fairness while using AI generated comics for blog posts…
- deleted 1y ago[deleted]
- positiveblue 1y ago> An allowlist run by one company that site owners chose to engage with. Exactly, no problem with that, just hinting that's not a protocol. > But the irony of taking an ideological stance about fairness while using AI generated comics for blog posts Wait, what?
- avtar 1y ago> Wait, what? I was referring to the following image: https://substackcdn.com/image/fetch/$s_!zRK-!,w_1250,h_703,c_fill,f_webp,q_auto:good,fl_progressive:steep,g_center/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F508ed5ad-3ced-4ae4-a881-b062861defc3_1024x1024.png https://substackcdn.com/image/fetch/$s_!zRK-!,w_1250,h_703,c...
- positiveblue 1y agoI know the image, what I do not understand is the argument between using it being incompatible with "fairness" and "openness"
- jaredcwhite 1y agoGenAI image generation is not fair. also: “Cloudelare” ;-P
- ryukoposting 1y agoI can't speak for the other commenter, but I think companies like Midjourney and OpenAI are robber barons exploiting people's creative work in ways that obviously aren't fair, but that our legal system wasn't equipped to prevent.
- pluto_modadic 1y agoI think it shouldn't require registering /with/ cloudflare. cloudflare should just look up the .well-known referenced and double check for impersonation, and keep score on how well behaved each one is.
- positiveblue 1y agoThis is one of the main points :+1:
- jeroenhd 1y agoUsing completely automated means would leave open the possibility to set up a new signature for every single request, or for batches of requests. The manual step is to cut down on the amount of automated abuse.
- mmaunder 1y agoBrought to you by substack. ;-) Seriously though, great post and a great conversation starter.
- positiveblue 1y agoI actually thought about this before publishing it hahaha Good thing they are not the only place to post!
- seanvelasco 1y agoas a Cloudflare customer, I am happy with their proposition. I personally do not want companies like Perplexity that fake their user-agent and ignore my robots.txt to trespass. and isn't this why people sign up with Cloudflare in the first place? for bot protection? to me, this is just the same, but with agents. i love the idea of an open internet, but this requires all party to be honest. a company like Perplexity that fakes their user-agent to get around blocks disrespects that idea. my attitude towards agents is positive. if a user used an LLM to access my websites and web apps, i'm all for it. but the LLM providers must disclose who they are - that they are OpenAI, Google, Meta, or the snake oil company Perplexity
- positiveblue 1y agoThe point is "should everyone just have an account with Cloudflare then"
- chrisweekly 1y agoYour complaints about "faking their user-agent" reminds me of this 15-year-old but still-relevant, classic post about the history of the user-agent string: https://webaim.org/blog/user-agent-string-history/ https://webaim.org/blog/user-agent-string-history/ TLDR the UA string has always been "faked", even in the scenarios you might think are most legitimate.
- jeroenhd 1y agoThe traditional UA fakery (adding Mozilla to the start and then just tacking on browser engine names) was the result of outdated websites breaking browsers. The problematic fakery here is that bots are pretending to be people by emulating browsers to prevent rate limits and other technical controls. That second category has also been with us since the dawn of the internet, but it has always been something worth complaining about. No trustworthy tool or service will pretend to be a real browser, at least not by default. If AI agents just identified themselves as such, we wouldn't need elaborate schemes to block them when they need to be blocked.
- CharlesW 1y agoDiscussion from yesterday: https://news.ycombinator.com/item?id=45055452 https://news.ycombinator.com/item?id=45055452
- hn_throw_250829 1y agoGood. Accelerate.
- imiric 1y ago> When I’m driving, I hand my phone to a friend and say, “Reply ‘on my way’ to my Mom.” They act on my behalf, through my identity, even though the software has no built-in concept of delegation. That is the world we are entering. That is a very small part of the world we're entering. The other vast majority of use cases will come from even more abusive bots than we have today, filling the internet with spam, disinformation, and garbage. The dead internet is no longer a theory, and the future we're building will make the internet for bots, by bots. Humans will retreat into niche corners of it, and those who wish to participate in the broader internet will either have to live with this, or abide by new government regulations that invade their privacy and undermine their security. So, yes, confirming human identity is the only path forward if we want to make the internet usable by humans, but I do agree that the ideal solution will not come from a single company, or a single government, for that matter. It will be a bumpy ride until we figure this out.
- dlcarrier 1y agoI use uncommon web browsers that don't leak a lot of information. To Cloudflare, I am indistingushable from a bot. Privacy cannot exist in an environment where the host gets to decide who access the web page. I'm okay with rate limiting or otherwise blocking activity that creates too much of a load, but trying to prevent automated access is impossible withou preventing access from real people.
- verdverm 1y agoThe website owner has rights too. Are you arguing they cannot choose to implement such gatekeeping to keep their site operating in a financially viable manner?
- SoftTalker 1y agoIf you put your information freely on the web, you should have minimal expectations on who uses it and how. If you want to make money from it, put up a paywall. If you want the best of both worlds, i.e. just post freely but make money from ads, or inserting hidden pixels to update some profile about me, well good luck. I'll choose whether I want to look at ads, or load tracking pixels, and my answer is no.
- rustc 1y ago> If you put your information freely on the web, you should have minimal expectations on who uses it and how. Does this only apply to "information" or should we treat all open source code as public domain?
- pessimizer 1y agoAll "open source" code was already pretty much public domain. All they'd have to do was put a page of OSI-approved licenses up on the site, right? An index of Open Source projects and their authors? Is this more than a weeks work to comply? Free Software is the only place where this is a real abridgement of rights and intention, and it's over. They've already been trained on all of it, and no judge will tell them to stop, and no congressman will tell them to stop.
- Wertulen 1y agoI suppose it’s time AI proselytization rediscovered the tragedy of the commons.
- ChrisArchitect 1y agoRelated: Web Bot Auth https://news.ycombinator.com/item?id=45055452 https://news.ycombinator.com/item?id=45055452 and associated blog post: The age of agents: cryptographically recognizing agent traffic https://blog.cloudflare.com/signed-agents/ https://blog.cloudflare.com/signed-agents/
- tzury 1y agoat its foundation, the bots issue is in fact 3 main issues: bots vs humans: humans are trying to buy tickets that were sold out to a bot data scrapping: you index my data (real estate listing) to not to route traffic to my site as people search for my product, as a search engine will do, rather to become my competitor. spam (and scam): digital pollution, or even worse, trying to input credit card, gift cards, passwords, etc. (obviously there are more, most which will fall into those categories, but those are the main ones) now, in the human assisted AI, the first issue is no longer an issue, since it is obvious that each of us, the internet users, will soon have an agent built into our browser. so we will all have the speedy automated select, click and checkout at our disposal. Prior to LLM era, there were search engines and academic research on the right side of the internet bots, and scrappers and north to that, on the wrong side of the map. but now we have legitimate human users extending their interaction with an LLM agent, and on top of it, we have new AI companies, larger and smaller which thrive for data in order to train their models. Cloudflare simply trying to make sense of this, whilst maintaining their bot protection relevant. I do not appreciate the post content whatsoever, since it lacks or consistency and maturity (a true understanding of how the internet works, rather than a naive one). when you talk about "the internet", what exactly are you referring to? a blog? a bank account management app? a retail website? social media? those are all part of the internet and each is a complete different type of operation. EDIT: I've written a few words about this back in January [1] and in fact suggested something similar: Leading CDNs, CAPTCHA providers, and AI vendors—think Cloudflare, Google reCAPTCHA, OpenAI, or Anthropic could collaborate to develop something akin to a “tokenized machine ID.” https://blog.tarab.ai/p/bot-management-reimagined-in-the https://blog.tarab.ai/p/bot-management-reimagined-in-the
- viktorcode 1y agoI wish Cloudflare would roll out AI poisoning attack as protection for their clients (providing bad data cache to AI bots), instead of this. Would work like a charm.
- meindnoch 1y agoThe private tracker community have long figured this out. Put content behind invite-only user registration, and treeban users if they ever break the rules.
- PaulRobinson 1y agoThis doesn't scale to the general web, does it? I think invite-only might work to build communities, but you end up in the situation we're in today where people are buying/selling invites, and that's with treebans in place. I do fear the actions of the current bot landscape is going to lead to almost everything going behind auth walls though, and perhaps even paid auth walls.
- lucb1e 1y agoI've been considering making this for the web. Why wouldn't it scale? Those selling invites would get banned soon enough if the people they distribute their invite to then send abusive traffic. Mystery shoppers can also make that a risky business if it's disallowed to sell invites (forcing them to be mostly free, such that the giver has nothing to gain from inviting someone who is willing to pay) One of the practical problems I rather saw was bootstrapping: how to convince any website owner to use it, when very few people are on the system? Where should they find someone to get invites from? As for tracking (auth walls), the website needs not know who you are. They just see random tokens with signatures and can verify the signature. If there's abuse, they send evidence to the tree system, where it could be handled similarly to HN: lots of flags from different systems will make an automated system kick in, but otherwise a person looks at the issue and decides whether to issue a warning or timeout. (Of course, the abuse reporting mechanism can also be abused so, again similar to HN, if you abuse the abuse mechanism then you don't count towards future reports.) Ideally, we'd not need this and let real judges do the job of convicting people of abuse and computer fraud, but until such time, I'd rather use the internet anonymously with whatever setup I like than face blocks regularly while doing nothing wrong
- PaulRobinson 1y agoI don't think it scales, because I'm not sure it scales on private trackers already. I'm not deep into that space, but I think there's a lot of problems with it that will scale as adoption scales, particularly around policing the sale of invites - the hope would be it self-police through treebanning, but I'm not sure it does. I think a sort of pseudo-anonymous auth system with backed in invites and treebans that website owners could easily adopt is interesting though. I'm not sure it's a business - for adoption reasons it likely needs to be a protocol - but it's an interesting idea, if it doesn't just turn into a huge admin headache for publishers.
- IshKebab 1y ago[flagged]
- thanatos_dem 1y agoAllowlist is arguably fitting for a list of things which are allowed.
- IshKebab 1y agoIt's called a whitelist. A perfectly good word that isn't racist and one that normal people are quite happy to use. As far as I can tell the allow/blocklist craze hasn't made it out of the software world.
- krapp 1y agoBoth whitelist and allowlist are equally normal and good. It's weird that people will claim that "politics" have no place in software while insisting that there is one and only one term "normal" people should use because the politics of the people who object to it are bad and wrong.
- zzo38computer 1y agoI agree that both words are good, but there is a difference. Whitelist means that anything explicitly listed (in the "whitelist" or "allow list") is allowed (or included, etc) and other stuff is disallowed (or excluded) by default (although in some cases, a program (or something else) might ask instead of forcibly blocking access). It is a compound word; you should not use a space or hyphen. (Using two words "white list" may be appropriate when you are refering to colours, e.g. the white list includes the list of whatever documents are to be copied on white paper, or "white list" might mean the list that is printed on white paper.) Allow list (I do not like the compound word; I think they should be separated and it looks better that way) is the list of what is allowed. (So, normally, this would mean that other stuff is not allowed, so it is still whitelisting.) In situations where colours would be involved and using words such as "whitelist" would be confusing, such words should be avoided, in order to avoid confusion.
- narrator 1y agoWait till these robots get out in the real world and start overwhelming real world resources.
- moribvndvs 1y agoI’m not necessarily coming to the defense of CF’s proposed solution, but it’s ridiculous and rather telling that the article mounts such a strong defense for agents around the notion they are simply completing user-directed tasks the user would otherwise do themselves, while avoiding the blatantly obvious issues of copyright, attribution, resource overusage, etc. presented by agents. It’s somewhat ironic to let fly the “free and open internet” battle cry on behalf of an industry that is openly destroying it.
- hoppp 1y agoShould use a public blockchain for this? Its good for it, store public keys, verify signatures etc.. none of that token stuff tho
- positiveblue 1y agoNo
- TYPE_FASTER 1y agoI think the reality is, we need identity on both the client and server sides. At some point soon, if not now, assume everything is generated by AI unless proven otherwise using a decentralized ID. Likewise, on the server side, assume it’s a bot unless proven otherwise using a decentralized ID. We can still have anonymity using decentralized IDs. An identity can be an anonymous identity, it’s not all (verified by some central official party) or nothing. It comes down to different levels of trust. Decoupling identity and trust is the next step.
- verdverm 1y agoDID spec, also used in ATProto, is quite flexible. It would be nice to see it used in more places and processes https://www.w3.org/TR/did-1.1/ https://www.w3.org/TR/did-1.1/
- lucb1e 1y agoIt's called an IP address. Since some ISPs don't assign a fixed IP to a subscriber, a timestamp is nowadays necessary. The combination is traceable to a subscriber who is responsible for the line, either to work with law enforcement if subpoenaed or to not send abusive traffic via the line themselves Why law enforcement doesn't do their job, resulting in people not bothering to report things anymore, is imo the real issue here. Third party identification services to replace a failing government branch is pretty ugly as a workaround, but perhaps less ugly than the commercial gatekeepers popping up today
- derefr 1y ago> Without that, I can simply hand the passport to another agent, and they can act as if they were me. This isn't the problem Cloudflare are trying to solve here. AI scraping bots are a trigger for them to discuss this, but this is actually just one instance of a much larger problem — one that Cloudflare have been trying to solve for a while now, and which ~all other cloud providers have been ignoring. My company runs a public data API. For QoS, we need to do things like blocking / rate-limiting traffic on a per-customer basis. This is usually easy enough — people send an API key with their request, and we can block or rate-limit on those. But some malicious (or misconfigured) systems, may sometimes just start blasting requests at our API without including an API key. We usually just want to block these systems "at the edge" — there's no point to even letting those requests hit our infra. But to do that, without affecting any of our legitimate users, we need to have some key by which to recognize these systems, and differentiate them from legitimate traffic. In the case where they're not sending an API key, that distinguishing key is normally the request's IP address / IP range / ASN. The problematic exception, then, is Workers/Lambda-type systems (a.k.a. Function-as-a-Service [FaaS] providers) — where all workloads of all users of these systems come from the same pool of shared IP addresses. --- And, to interrupt myself for a moment, in case the analogy isn't clear: centralized LLM-service web-browsing/tool-use backends, and centralized "agent" orchestrators, are both effectively just FaaS systems, in terms of how the web/MCP requests they originate, relate to their direct inbound customers and/or registered "agent" workloads. Every problem of bucketing traditional FaaS outbound traffic, also applies to FaaSes where the "function" in question happens to be an LLM inference process. "Agents" have made this concern more urgent/salient to increasingly-smaller parts of the ecosystem, who weren't previously considering themselves to be "data API providers." But you can actually forget about AI, and focus on just solving the problem for the more-general category of FaaS hosts — and any solution you come up with, will also be a solution applicable to the "agent formulation" of the problem. --- Back to the problem itself: The naive approach would be to block the entire FaaS's IP range the first time we see an attack coming from it. (And maybe some API providers can get away with that.) But as long as we have at least one legitimate customer whose infrastructure has been designed around legitimate use of that FaaS to send requests to us, then we can't just block that entire FaaS's IP range. (And sure, we could block these IP ranges by default, and then try to get such FaaS-using customers to send some additional distinguishing header in their requests to us, that would take priority over the FaaS-IP-range block... but getting a client engineer to implement an implementation-level change to their stack, by describing the needed change in a support ticket as a resolution to their problem, is often an extreme uphill battle. Better to find a way around needing to do it.) So we really want/need some non-customer-controlled request metadata to match on, to block these bad FaaS workloads. Ideally, metadata that comes from the FaaS itself. As it turns out, CF Workers itself already provides such a signal. Each outbound subrequest from a Worker gets forcibly annotated "on the way out" with a request header naming the Worker it came from. We can block on / rate-limit by this header. Works great! But other FaaS providers do not provide anything similar. For example, it's currently impossible to determine which AWS Lambda customer is making requests to our API, unless that customer specifically deigns to attach some identifying info to their requests. (I actually reported this as a security bug to the Lambda team, over three years ago now.) --- So, the point of an infrastructure-level-enforced public-visible workload-identity system, like what CF is proposing for their "signed agents", isn't just about being able to whitelist "good bots." It's also about having some differentiable key that can cleanly bucket bot traffic, where any given bucket then contains purely legitimate or purely malicious/misbehaving bot traffic; so that if you set up rate-limiting, greylisting, or heuristic blocking by this distinguishing key, then the heuristic you use will ensure that your legitimate (bot) users never get punished, while your misbehaving/malicious (bot) users automatically trip the heuristic. Which means you never need to actually hunt through logs and manually blacklist specific malicious/misbehaving (bot) users. If you look at this proposal as an extension/enhancement of what CF has already been doing for years with Workers subrequest originating-identity annotation, the additional thing that the "signed agents" would give the ecosystem on behalf of an adopting FaaS, is an assurance that random other bots not running on one of these FaaS platforms, can't masquerade as your bot (in order to take advantage of your preferential rate-limiting tier; or round-robin your and many others' identities to avoid such rate-limiting; or even to DoS-attack you by flooding requests that end up attributed to you.) Which is nice, certainly. It means that you don't have to first check that the traffic you're looking at originated from one of the trustworthy FaaS providers, before checking / trusting the workload-identity request header as a distinguishing key. But in the end, that's a minor gain, compared to just having any standard at all — that other FaaSes would sign on to support — that would require them to emit a workload-identity header on outbound requests. The rest can be handled just by consuming+parsing the published IP-ranges JSON files from FaaS providers (something our API backend already does for CF in particular.)
- jbrisson 1y ago"In the 90s, Microsoft tried to “embrace and extend” the web, but failed. And that failure was a blessing." Basically MS tried to kill the web with their Win95 release, the infamous Internet Explorer and their shitty IIS/Frontpage tandem. I deeply hate them since that day.
- positiveblue 1y agomany people don't remember/know history though
- fortran77 1y agoCloudflare lost a lot of credibility by backing off its "neutral" stance and booting certain sites--some which were admittedly horrible--from the their service. Now it seems they want to be even more of a gatekeeper.
- ramoz 1y agoWe dont need gatekeepers. We do need to verify agents that act, in a reasonable way, on behalf of human vs an agent swarm/bot-mining operation (whether conducted by a large lab or a kid programming claude code to ddos his buddy's next.js deployment).
- madrox 1y agoThe web doesn't need gatekeepers the way you don't need a bank account, driver's license, or a credit card. You can do without it, but it sure makes it harder to interact with modern society. The days of the mainstream internet being a libertarian frontier are more or less over. The capitalist internet is firmly in charge. The real question is whether there is more business opportunity in supporting "unsigned" agents than signed ones. My hope is that the industry rejects this because there's more money to be made in catering to agents than blocking them. This move is mostly to create a moat for legacy business. Also, if agents do become the de-facto way of browsing the internet, I'm not a fan of more ways of being tracked for ads and more ways for censorship groups to have leverage. But the author is making a strawman argument over a "steelman" argument against signed agents. The strongest argument I can see is not that we don't need gatekeepers, but that regulation is anti-business.
- jaredcwhite 1y agoThis article can easily be dismissed when hardly a moment in you see the headline "Agents Are Inevitable" I'm sorry, but the "agents" of "agentic AI" is completely different from the original purpose of the World-Wide Web which was to support user agents. User agents are used directly by users—aka browsers. API access came later, but even then it was often directed by user activity…and otherwise quite normally rate-limited or paywalled. The idea that now every web server must comply with servicing an insane number of automated bots doing god-knows-what without users even understanding what's happening a lot of the time, or without the consent of content owners to have all their IP scraped into massive training datasets is, well, asinine. That's not the web we built, that's not the web we signed up for; and yes, we will take drastic measures to block your ass.
- 1gn15 1y agoSpeak for yourself. This is just the semantic web: a web not built just for humans, but also for robots or any other types of agents that may wish to build upon the data. User agents never meant just web browsers, and operators blocking based on it necessitated hiding your identity. Blocking bots is an absurd and unwinnable proposition, just like DRM; there's always the final, nuclear option of the analog hole, a literal video camera pointed at a monitor and using a keyboard and mouse. If you really need to, deploy a proof of work shield that doesn't discriminate against user agents, just like what Onionsites do.
- hugs 1y agoor we could just require postage. (HTTP status code 402) one potential solution: https://www.l402.org/ https://www.l402.org/
- ctoth 1y agoThe web doesn't need attestation. It doesn't need signed agents. It doesn't need Cloudflare deciding who's a "real" user agent. It needs people to remember that "public" means PUBLIC and implement basic damn rate limiting if they can't handle the traffic. The web doesn't need to know if you're a human, a bot, or a dog. It just needs to serve bytes to whoever asks, within reasonable resource constraints. That's it. That's the open web. You'll miss it when it's gone.
- johncolanduoni 1y agoBasic damn rate limiting is pretty damn exploitable. Even ignoring botnets (which is impossible), usefully rate limiting IPv6 is anything but basic. If you just pick some prefix from /48 to /64 to key your rate limits on, you'll either be exploitable by IPs from providers that hand out /48s like candy or you'll bucket a ton of mobile users together for a single rate limit.
- ctoth 1y agoYou make unauthenticated requests cheap enough that you don't care about volume. Reserve rate limiting for authenticated users where you have real identity. The open web survives by being genuinely free to serve, not by trying to guess who's "real." A basic Varnish setup should get you most of the way there, no agent signing required!
- hombre_fatal 1y agoYour response to unauthenticated requests could be <h1>Hello world</h1> served from memory and your server/link will still fail under a volumetric attack, and you still get the pleasure of paying for the bandwidth. So no, this advice has been outdated for decades. Also you're doing some sort of victim blaming where everyone on earth has to engineer their service to withstand DoS instead of outsourcing that to someone else. Abusers outsource their attacks to everyone else's machine (decentralization ftw!), but victims can't outsource their defense because centralization goes against your ideals. At least lament the naive infrastructure of the internet or something, sheesh.
- calmbonsai 1y agoWhile I concur with the effective tech, I don't think this is something that's a net win for society. Just because you can, doesn't mean you should and I don't feel any one entity (private or public) should be an arbiter on these matters. This is something that can, and should, be negotiated at the "last virtual mile".
- positiveblue 1y ago100% needs to be done at the last mile.
- rs_rs_rs_rs_rs 1y ago>Just because you can, doesn't mean you should and I don't feel any one entity (private or public) should be an arbiter on these matters. What do you mean by private? Should I not be allowed to block AI agents on my sites using Cloudflare?
- calmbonsai 1y agoMore precisely, Cloudflare should not be able to offer this service. AI agents should be blocked at the hosted API endpoint ("last mile"), not at a CDN or other sort of intermediary. That said, if you're using Cloudflare Workers, where the endpoint itself is hosted by Cloudflare, that would be ethical.
- kentonv 1y agoIf the CDN is implementing the explicitly-configured preferences of the specific site, what difference does it make if the blocking happens at the CDN vs. the site's origin server?
- calmbonsai 1y agoIf it's via a CDN pull, that's totally cool. IF it's with a push or acting as a pure MIM, that's not cool.
- zzo38computer 1y agoWith what they say about authorization, I think X.509 would help. (Although central certificate authorities are often used with X.509, it does not have to be that way; the service you are operating can issue the certificate to you instead, or they can accept a self-signed certificate which is associated with you the first time it is used to create an account on their service.) You can use the admin certificate issued to you, to issue a certificate to the agent which will contain an extension limiting what it can be used for (and might also expire in a few hours, and also might be revoked later). This certificate can be used to issue an even more restricted certificate to sub-agents. This is already possible (and would be better than the "fine-grained personal access tokens" that GitHub uses), but does not seem to be commonly implemented. It also improves security in other ways. So, it can be done in such a way that Cloudflare does not need to issue authorization to you, or necessarily to be involved at all. Google does not need to be involved either. However, that is only for things where would should normally require authorization to do anyways. Reading public data is not something that should requires authorization to do; the problem with this is excessive scraping (there seems to be too many LLM scraping and others which is too excessive) and excessive blocking (e.g. someone using a different web browser, or curl to download one file, or even someone using a common browser and configuration but something strange unexpected happens, etc); the above is something unrelated to that, so certificates and stuff like that does not help, because it solves a different problem.
- jeroenhd 1y agoWhat problem does this solve that a basic API key doesn't solve already? The issue with that approach is that you will require accounts/keys/certificates for all hosts you intend to visit, and malicious bots can create as many accounts as they need. You're just adding a registration step to the crawling process. Your suggested approach works for websites that want to offer AI access as a service to their customers, but the problem Cloudflare is trying to solve is that most AI bots are doing things that website owners don't want them to do. The goal is to identify and block bad actors, not to make things easier for good actors. Using mTLS/client certificates also exposes people (that don't use AI bots) to the awful UI that browsers have for this kind of authentication. We'll need to get that sorted before an X509-based solution makes any sense.
- hinkley 1y agoI used to joke that I worked for the last DotCom startup, a company that got a funding round after the shit hit the fan. They were working on an idea that looked a bit like an RSS feed for an entire website, where you would run your own spider and then our search engine could hit an endpoint to get a delta instead of having to scan your entire site. If they’d made the protocol open instead of proprietary, we maybe could have gotten spiders to play nicer since each spider after the first would be cheaper, and eventually maybe someone could build pub sub hooks into common web frameworks to potentially skip the scan entirely for read-mostly websites, generating delta data when your data changed. But of course when the next round of funding came due nobody was buying. I thought about this a lot on my last project, where spiders were our customers’ biggest users. One of those apps where customer interactions were intense but brief and the rank in Google mattered equally with all other concerns. Nobody had architected for the actual read/write workflow of the system of course, and that company sold to a competitor after I left. Who migrated all customers to their solution and EOLed ours for being too fat in a down economy.
- JohnMakin 1y ago> The same is true online. A cryptographic signature that claims “I am acting on behalf of X” means nothing unless it is tied to something real, like a verifiable infrastructure or a range of IPs. Without that, I can simply hand the passport to another agent, and they can act as if they were me. The passport becomes nothing more than a token anyone can pass around. how does this person think jwt’s work?
- positiveblue 1y agoHi, "this person here" Cloudflare will block that request that has a jwt because "it does not come from a person". What I was trying to say is that even the discussion "is this a bot 100% sure or not" makes no sense.
- Animats 1y agoAre bots using a large number of IP addresses simultaneously, so they look like a DDOS attack? Or are they just making ordinary requests from a small number of addresses. If it's the latter, all you need is some kind of fair queuing so those requests compete with each other for access, not with other users.
- JohnMakin 1y agoOften it is rotating residential proxies. It is virtually impossible to mitigate this behavior from the IP level.
- fooey 1y agoThey're using state of the art obfuscation that makes them indistinguishable from malicious botnets. It's an arms race with billion dollar companies vying to consume the most content before it all collapses the open web is dead and whatever's left will be locked be authentication and paywalls
- jeroenhd 1y agoBots are probing for access from various servers, eventually falling back to executing requests from residential IP addresses: https://blog.cloudflare.com/perplexity-is-using-stealth-undeclared-crawlers-to-evade-website-no-crawl-directives/ https://blog.cloudflare.com/perplexity-is-using-stealth-unde... Cloudflare is dealing with a couple million faked requests every day just from Perplexity users, and Perplexity is far from the worst player in the field. The problem would be quite easy to solve with basic rate limiting if it weren't for the attempts to bypass access controls.
- sys_64738 1y agoCloudflare slows the whole damn websites down. It takes many seconds to deal with their trash. I hope they crash and burn. Let's get back to very low latency websites without the cloudflare garbage.
- const_cast 1y agoCloudflare as a CDN greatly greatly speeds up the web. All the custom code they write on top of that to transform HTML for you? Ehhhh... don't use those features. Most are easily reproducible on the backend.
- p3rls 1y ago[flagged]
- timpera 1y agoCloudflare is being really annoying lately. It looks like they desesperately want to close the web to get their 30% fee on AI crawling fees.
- neximo64 1y agoSo Cloudflare becomes the gatekeeper then? I kind of want my site to be indexed with agents and used without any interference
- jeroenhd 1y agoBy not using Cloudflare your website will be indexed by everyone. The gatekeeper aspect only applies if you use Cloudflare to distribute your website (and even then Cloudflare offers options to control this bot shield thing).
- neximo64 1y agoI want it to be indexed by everyone, thats the whole point. So what then Cloudflare can use all these websites as leverage against Google, OpenAI and Microsoft? I kind of want my content to be indexed.
- jeroenhd 1y agoThe content you host will only be blocked from being indexed if you decide to use a service that blocks indexing. If you host your content on other people's services, then you never had the power to make that decision anyway. If you want your content to be indexed, simply don't use Cloudflare. Host your own servers. Use a different CDN if you want the benefits of Cloudflare's networks.
- ymyms 1y agoI agree with pretty much everything the author has said. I’ve been looking at the problem more on the enterprise side of things: how do you control what agents can and can’t do on a complex private network, let alone the internet. I’ve actually just built an “identity token” using biscuit that you can delegate however you want after. So I can authenticate (to my service, but it could be federated or something just as well), get a token, then choose to create a delegated identity token from that for my agent. Then my agent could do the same for subagents. In my system, you then have to exchange your identity token for an authorization token to do anything (single scope, single use). For the internet, I’ve wondered about exchanging the identity token + a small payment (like a minuscule crypto amount) for an authorization token. Human users would barely spend anything. Bots crawling the web would spend a lot.
- IAmGraydon 1y agoI do like Cloudflare in general, but the whole anti-AI push is just another form of the Luddism surrounding AI since 2022. Cloudflare perhaps wisely picked up on this trend and decided to capitalize on it, but I think it would be a mistake to allow it to become their brand.
- djoldman 1y agoSorry, the "web" isn't "open" and hasn't been for a while. Most interaction, publication, and dissemination takes place behind authentication: Most social media, newspapers, etc. throttle, block, or otherwise truncate non-authenticated clients. Blogs are an extremely small tranche of information that the average netizen consumes.
- coldtea 1y ago>Sorry, the "web" isn't "open" and hasn't been for a while. Doesn't matter, it still doesn't need gatekeepers, and if they're already a lot, it should reduce them, not increase them.
- jrochkind1 1y ago> The same is true online. A cryptographic signature that claims “I am acting on behalf of X” means nothing unless it is tied to something real, like a verifiable infrastructure or a range of IPs. Without that, I can simply hand the passport to another agent, and they can act as if they were me. The passport becomes nothing more than a token anyone can pass around. Well, that's true of any crytpographic key? In this case, it would mean you are giving them permission to act on your behalf. Nothing wrong with that. If some of the people acting on your behalf start acting maliciously, then presumably those who decided to trust the people who were acting on your behalf would stop doing so. Is this not common to how most any digital authentication works at all? You can always share your keys. That's a feature not a bug, when the actor you want to identify is meant to have a distributed implementation. I understand the concern about how much power CloudFlare has, how they have the ability to gatekeep a large part of the internet. Absolutely, this is alarming. But the Web Both Auth protocol itself is not the problem -- it seems to me to be written and designed appropriately for authentication of automated web agents. And I think we desperately need something for that. I, like many people, are being forced to put bot precautions in place, because otherwise my sites are overwhelmed. But this means I wind up blocking bots that I don't want to block too. Because they are are partners, because I approve of what they are doing, becuase they have demonstrated good behavior. I have no way to do that right now. IP address ranges are absolutely not the right way. IP addresses are network topology, not authentication. i worked in academia for some time, where large unviersities have a history of trying to use IP addresses for authentication -- and even working with internal IP addresses theoretically controlled by the (large) institution, it was a fool's game. IP addresses can change all the time -- even for a device which has not moved it's physical location. Plus resources can be allocated to different physical locations. Different actors can share an IP address. They are often changed at various lower levels of hiearchical administration without informing the top, for network topological concerns -- they are designed for this. Etc etc etc. I understand the concern about CloudFlare's gatekeeping monopoly. There may be ways that Web Both Auth can make it worse. Discussion of that is not inappropriate. Maybe there are ways to ameliorate it (will individual customers be ablet o have their own allow-lists? Can we insist on that? Is that enough?). Maybe not good enough. But let's focus the discussion on that -- there is in fact nothing wrong with Web Both Auth protocol, at least nothing covered in this essay, it is well-designed for authenticating bot agents, and we actually do need something that does that, in the current world where misbehaving disguised bot agents have become a real problem. Not having a way to authenticate distributed bot actors who wish to opt in to a way to be authenticated (everyone else is free to try to evade the bot detectors same as they are now?) -- is going to create more damage. All these people railing against what seems to be an appropriate protocol for authentication because they don't like Cloudflare's monopoly are distressing me, it's going to be worse if we don't have a way to do it. It is an open protocol not just for use by cloudflare.
- account42 1y ago> Yes, identity for agents is a real problem. I don't agree that bot identity is a problem or something that we are better off with if it is solved. I'd rather have a web where adversarial interoperability is possible than one where service operators have a say what tools you can use to access their websites. The main problem with we are see today is bad actors that a) completely ignore and/or side step copyright an licensing to use work of others for their own benefit without contributing anything back b) send an unreasonable number of request Both will not be solved with identifying bots, that will at best get rid of the small players and give Google, Meta, etc. even more power. The unsustainable parasitic theft of open content is something that needs to be dealt with legally and nothing else will solve it. DRM never works. If enough people block Gemini, Google will just feed it with Google bot crawls and no one can afford to block that. And then they will sell the data to other players. Or someone will make a browser extension to do the same. The second issue also should be solved via legislation and enforcement thereof. It can also be solved by disconnecting and/or throttling abusive networks wholesale - whole countries if need be. You know, like we have been handling abusive network participants forever. Trying to detect "bots" is a fools errand that will only get rid of the laziest crawlers. You cannot win the bot blocking game when the bots can afford to spend more resources per request than real users are willing to. So yes, the web must remain open. But to do that we must not have bot identity checks at all - whether that's managed by a single company that has inserted it as a gatekeeper or not.