6 ms·
Stay discoverable in search while disallowing AI training
- AnonC 17d ago> Accountable mixed-use crawlers remain allowed for search. Every other training crawler is blocked, including the training-only crawlers run by Amazon, Anthropic, Meta, and OpenAI — blocking those does not affect search. > We also categorize the relevant crawlers from Amazon, Anthropic, Meta, and OpenAI as Accountable. These organizations separate their Search and Training crawlers, so Cloudflare can block the Training crawler without affecting search. I find it difficult to trust that either Meta or OpenAI would use their separate search and training crawlers only for the respective purposes. Their pinky promises have no value, IMO. Both companies are premised on deceptive behaviors.
- gchamonlive 17d agoCloudflare, enabling the problem and the solution since, how long has it been?
- arm32 17d ago(checks watch) 17 years Fun lava lamp story, though
- gleezard 17d agoA bit too late honestly (?). With so many people who have shifted over to reading AI summaries as a primary search response, those with AI-enabled sites will win by attrition. There is no going back from this. And the internet is a relatively new phenomenon. Recklessly, blindly applying ads to pages in hopes of generating revenue is a very silly thing to do. Technology with ad blockers and now AI summaries has taken that away. New business models, perhaps actually decent ones are required. Death of ads everywhere? Good fucking riddance. Posted from LibreWolf.
- tracerbulletx 17d agoAll this attitude does is tear down the only viable income source for independent publishers and demonizes them for trying to make money, while everyone let's huge corporations off the hook for it because "well that's just what they do"
- gleezard 17d agoAdvertising in the way it’s done is demonic in and of itself. I don’t care - find a better business model.
- antonvs 16d agoA better business model won’t help. What you need is a better economic and political system. If you “let the market decide” what advertising looks like, you get what we have today, because “let the market decide” is an incoherent claim that’s really shorthand for the rule of a wealthy minority. Advertising is just a convenient way to extract wealth from a society without having to go to all the trouble of satisfying customers.
- octoberfranklin 17d agoIndependent publishers can't afford to run an ad network. You aren't independent. You work for the BigTech company that serves ads on your site.
- lostmsu 16d agoPapers charged per "user" since times immemorial.
- arm32 17d agoWhat does this new setting actually do? Does it block their IP ranges too, I hope? As if Meta, for instance, is actually going to respect Accountable, via themselves or their partners, quite frankly is eyebrow raising at best.
- ankurshv 17d ago[flagged]
- zergrush 17d agowhat weirds me out is the analytics theres no way my index.html page with nothing is getting 10000 hits a day wtf?
- itake 17d ago7 per minute. I use my high school’s website to test Internet connectivity bc the domain is short and they don’t do a TLS redirect (making it easy to detect WiFi portals).
- JoshTriplett 17d agoneverssl.com
- DaSHacka 17d agoWish the admin never added the SSL redirect, its literally the namesake lol I just want basically captive.apple.com with a shorter domain and the webserver not even listening on port 443 at all. Surprised someone hasn't made this yet, it only requires one spare public IP.
- itake 17d agoI use this too, but my high school url is 8 characters. neverssl.com is 12. :)
- jesterson 16d agoThey bullshit big time on that, I have caught them not once. For a site with proven visitors 100,000 per month CF tell me they saved 90,000 over 1,000,000 visitors (rough numbers). Sure they know large numbers make people feel good.
- skybrian 17d agoI didn't know websites could opt out of providing data to Google's AI training. Looks Google added support for this via 'Google-Extended' in robots.txt back in 2023: https://blog.google/innovation-and-ai/products/an-update-on-web-publisher-controls/ https://blog.google/innovation-and-ai/products/an-update-on-...
- qtqtqt 16d agodoc: https://developers.google.com/crawling/docs/crawlers-fetchers/google-common-crawlers#google-extended https://developers.google.com/crawling/docs/crawlers-fetcher...
- mskalski 17d agoI wonder if protocols like Web Bot Auth [1] will see wider adoption. At least as a supported mechanism for those bots which identify themselves. The rest probably still have to be treated with Anubis. In my free time I've recently been experimenting with a Web Bot Auth implementation as an Envoy dynamic module [2] to have a way to define some additional policies for the traffic from bots. [1] https://datatracker.ietf.org/doc/draft-ietf-webbotauth-httpsig-protocol/ https://datatracker.ietf.org/doc/draft-ietf-webbotauth-https... [2] https://github.com/michalskalski/envoy-web-bot-auth https://github.com/michalskalski/envoy-web-bot-auth
- kinduff 17d agoI maintain a cloud IP ranges database, and I'm going to test this out. I have my doubts, though. A formal title like "Accountable" (capitalized) sounds deliberate, but I can't help imagining the renewal email: "Hey, want to renew your Accountable™ license? Just pinky promise again that you use your IPs for what you say you do."
- India_InfraNote 17d ago[flagged]
- nirmeetimthebes 17d ago"Accountable" is just a fancy word for "pinky promise, but with a label." Nothing stops the data from ending up in a training run once it's already been fetched.
- userbinator 17d ago...and the only way to stop[1] that is by effectively DRM'ing everything, which is a level of dystopia that I don't think even Stallman ever anticipated, nor do I want to happen. [1] Analog hole and other workarounds aside, naturally.
- jareklupinski 16d ago> "pinky promise, but with a label." and that label is "pending litigation"
- 1vuio0pswjnm7 17d ago"Cloudflare classifies bots by behavior, and a single bot can exhibit more than one behavior." Is that really true CF classifies anyone not using a popular browser with Javascript enabled as a "bot" CF fingerprints www users As an example, look at CF's Permissions-Policy HTTP response header on a site with CF "bot protection", i.e., the "checking your browser" CAPTCHA nonsense (challenges.cloudflare.com). Then look at IA's Permissions-Policy response header. One CDN is advertiser-focused, the other is user-focused IA = Internet Archive
- devmor 17d agoAs far as I can tell, after months of fighting being DDoSed by Anthropic and OpenAI across 50+ sites - Cloudflare also allows what it considers "good bots" through all of your bot blocking rules, with no option to turn this off unless you pay them money.
- RobotToaster 16d agoThe "good bots" also just happen to be from companies that pay cloudflare a lot of money, I imagine.
- actionfromafar 16d agoGoodness Tokens
- basilikum 16d agoHow did you dertmine it was those two companies? Also did you disallow them in robots.txt?
- devmor 16d agoNo, I definitely spent months fighting off bots across multiple hosts and never set up a robots.txt file anywhere nor did I look at the analytics dashboard on cloudflare.
- aaron695 17d ago[dead]
- dzhiurgis 17d agoIf this admin is serious about AI growth they’d make anti-scrapping illegal.
- DaSHacka 17d agoI don't think this admin even knows what scraping is. Unless it's "scraping the bottom of the barrel", something they're very familiar with.
- userbinator 17d agoJust make user-agent-discrimination illegal.
- GetSMS 17d agoThis seems useful for my site. I don't really want competitors' AI systems use our research data to train their models.
- gdiamos 17d agoI wouldn’t trust an AI company to honor this as far as I could throw them
- DharmaPolice 17d agoI feel like most of these schemes to categorise data as "public but not really" are ultimately doomed to failure. Even if you could trust every AI company in the world to respect these terms, is there anything stopping someone else indexing the data and selling them the information? I know there's copyright law but they're apparently ignoring that anyway. Ultimately this reminds me of those really early social media profiles (before people understood privacy settings if they even existed) which would say "If you're not my friend you're not allowed to read this page". If you don't want your content to end up in some database/archive don't publish it for the whole world to see.
- sillyfluke 16d agoIt's also a way to literally advertise to AI companies that you have some data worth plundering in an increasingly dead and sloppy internet that has diminishing returns for training.
- kixiQu 16d ago> If you don't want your content to end up in some database/archive don't publish it for the whole world to see. This principle somewhat reminds me of the line that "If you're not paying for the product, you are the product", and it seems to me similarly misleading - my data gets harvested and sold by companies with which I have non-paying relationships and by companies I have to pay for things (I am made the product in both cases). As you note - the AI companies are ignoring copyright law and pirating everything that seems useful to them regardless of whether it was published for free access. The potential externalities here are troubling. https://vbuckenham.com/blog/how-to-find-things-online/ https://vbuckenham.com/blog/how-to-find-things-online/
- qsbuilder 16d agoThe irony is that search engines are AI companies now. Telling them 'index me for search but don't train your models' is asking them to split a brain that’s already fully merged.
- tchalla 16d agoExactly. I don’t get this. You may decide to not train but your content will still show up in search.
- VBprogrammer 16d agoEven looking for companies which supply services is now far better on AI chats than Google. For me being visible in AI training is going to be more important than search in the next year or two. If I was a big AI company I'd certainly be tempted to make sure that anyone who excluded themselves from "AI training" also got themselves excluded from AI results.
- nullbio 16d agoWhat does this really mean though? You can use an LLM to search.
- nicolodev 16d agothere have been a few links here on hn about content protection based on Markov’s chain. It might be interesting for Cloudflare to add damage for the scraper that tries to query the page, and not just blocking them.
- fwlr 16d ago“Pre-label your most valuable data for us and we promise to not make it too obvious that we’re training on it”
- shark1 16d ago"Apple, Google, and Microsoft honor or have committed (in a specified time frame) to honor this setting." Specified Time Frame ;)
- RobotToaster 16d ago"infinity is a timeframe!"
- OroPla 16d agoIn a world where people's searches are already mostly answered by AI, is there any point to disallow the training? I could totally see a future where Google just stops showing you links to actual web pages altogether and just gives you their chatbot.
- deleted 16d ago[deleted]
- 833dong 16d ago[flagged]