118 ms·
I don't think it should. If a user asks the AI to read the web for them, it should read the web for them. This isn't a vacuum charged with crawling the web, it'
by _2d30 2y ago
I don't think it should. If a user asks the AI to read the web for them, it should read the web for them. This isn't a vacuum charged with crawling the web, it's an adhoc GET request.
- deleted 2y ago[deleted]
- internetter 2y agoYou could make this justification for a lot of unapproved bot activity.
- taskforcegemini 2y agoyou could, but this article is about claude.
- bayindirh 2y agoHow can you be so sure? Processors love locality, so they fetch the data around the requested address. Intel even used to give names to that. So, similarly, LLM companies can see this as a signal to crawl to whole site to add to their training sets and learn from it, if the same URL is hit for a couple of times in a relatively short time period.
- mvdtnz 2y agoNo thank you, when I define a robots.txt file I expect all automated systems to respect it.
- beeflet 2y agoSomeone should call the robots.txt police then, there's a bandit on the loose!
- TheDudeMan 2y agoBut this isn't automated. This is user-driven.
- jrflowers 2y agoIf this feature isn’t already part of the Claude API it likely will be at some point, in which case many Claude requests will be automated with no way to distinguish between user-driven or otherwise.
- pixl97 2y agoSimply put, at the end of the day you lose, AI blocking will not work. I mean, currently the AI request comes from the datacenter running the AI, but eventually one of two things will happen. AI models will get small/fast enough to run on user hardware and use the users resources: End result? You lose. The user will set their own headers and sites will play the impossible game of identifying AI. AI sites will figure out how to route the requests via any number of potential methods so the requests appear to come from the user anyway: End result? You lose. The sites attempting to block will play the cat and mouse game of figuring out what is AI or not AI. Note, this doesn't mean AI blocking isn't worth doing, if nothing else to reduce load on the servers. It's just not a long term winning strategy.
- ipaddr 2y agoWelcome to the world of CAPTCHAs
- pixl97 2y agoHeh, I'm always reminded of dunkey when captchas are brought up. It seems AI gets better faster at them than humans do. https://www.youtube.com/watch?v=WqnXp6Saa8Y https://www.youtube.com/watch?v=WqnXp6Saa8Y
- jimbokun 2y agoDepends if the legal system survives. You may not be able to stop AIs from crawling web sites through technological means. But you can confiscate all the resources of the company that owns the AI.
- victorbjorklund 2y agoA browser is automated too.
- goatlover 2y agoBrowser don't automatically browse, unless they are being automated.
- Sargos 2y agoAny AI tool I make will ignore robots.txt on principle. Artificial humans should have equal rights as real humans.
- navigate8310 2y agoThink of the "searching" LLM as a peon of the user, the user asks, the peon performs. In that essence, searching by the LLM should be human-driven and must not be blocked. It's just an automated system doing the search not your personal peon.
- bcrosby95 2y agoCan't you make the same argument for a crawler? The user wants information, the peon (crawler) just compiles it for them.
- theshackleford 2y agoThen you’ve fundamentally misunderstood what a robots.txt file does or is even intended to do and should reevaluate if you should be in charge of how access is or is not prevented to such systems. Absolutely nothing has to obey robots.txt. It’s a politeness guideline for crawlers, not a rule, and anyone expecting bots to universally respect it is misunderstanding its purpose.
- usrbinbash 2y ago> Absolutely nothing has to obey robots.txt And absolutely no one needs to reply to every random request from an unknown source. robots.txt is the POLITE way of telling a crawler, or other automated system, to get lost. And as is so often the case, there is a much less polite way to do that, which is to block them. So, the way I see it, crawlers and other automated systems have 2 options: They can honor the polite way of doing things, or they can get their packets dropped by the firewall.
- 1shooner 2y ago>You can now use Claude to search the internet to provide more up-to-date and relevant responses. It's a search engine. You 'ask it to read the web' just like you asked Google to, except Google used to actually give the website traffic. I appreciate the concept of an AI User-agent, but without a business model that pays for the content creation, this is just going to lead to the death of anonymously accessible content.
- darepublic 2y agoWell I expect eventually the agent will be able to act on your behalf with your credentials.
- elefanten 2y agoAnd as advertisers get declining human views on their ads, the value of the business model will dwindle until it needs to be replaced by other forms of revenue. Content that can't shift business models and requires revenue to continue will die off. Edit: Maybe that's fine, maybe that's bad. Maybe new models will emerge and things will reshape. But I'm just supporting the case that AI agents will pressure the current "free" content economy.
- beeflet 2y agothe free content economy is bogus, I am part of a growing segment of users that just block ads anyways.
- disiplus 2y agoI'm also and I pay for the services that I use to not see ads, but I don't pay for every single one. For example a local classified website is financed by ads, and I don't think anybody will pay for just looking at stuff there. Maybe they can switch to the model where the person puting the thing for sale would pay but hat is something where we are not currently.
- usrbinbash 2y ago> This isn't a vacuum charged with crawling the web, it's an adhoc GET request. Doesn't matter. The robots-exclusion-standard is not just about webcrawlers. A `robots.txt` can list arbitrary UserAgents. Of course, an AI with automated websearch could ignore that, as can webcrawlers. If they chose do that, then at some point, some server admins might, (again, same as with non-compliant webcrawlers), use more drastic measures to reduce the load, by simply blocking these accesses. For that reason alone, it will pay off to comply with established standards in the long run.
- renewiltord 2y agoIn the limit of the arms race it's sufficient for the robot to use the user's local environment to do the browsing. At that point you can't distinguish the human from the robot.
- usrbinbash 2y agoThat's not how many of these services work though. The websearch and subsequent analysis of the results by an LLM are done from the servers of whoever supplies the solution.
- scoofy 2y agoMany if not most websites are paid for by eyeballs not by get requests. A bot is a bot is a bot. Respect robots.txt or expect to have your IPs banned.
- danenania 2y agoIt may not be very long before the big majority of web searches are via AI. If that happens, blocking AI will mean blocking most people too. You’d already be blocking me as I’d guess I now search via AI >90% of the time between perplexity, chatgpt, deep research, and google search AI.
- scoofy 2y ago>It may not be very long before the big majority of web searches are via AI. If that happens, blocking AI will mean blocking most people too. If that happens a big majority of websites will go bankrupt and won't exist anymore to be searched. Problem solved!
- wraptile 2y agobig doubt on that and maybe that's a good thing? Let's be honest, right now most of the web is dominated by low effort spam. Taking money away from view farming would dramatically increase the web quality of the web. Suddently that guy who's really into "key gardening" doing research and publishing detailed results on his website actually has viewers — isn't this good? Especially since website hosting is close to being free these days.
- moooo99 2y ago> big doubt on that and maybe that's a good thing? Let's be honest, right now most of the web is dominated by low effort spam. I think that is funny considering it is likely going to have the exact opposite effect. Low effort blog spam is cheap to make. And it is often part of content marketing strategies where brand visibility is all that matters, so not much harm if the viability is directly on your site or in an AI chatbit interface. Quality content on the other hand is hard to make. And there are two groups of people who make such content: 1. individuals or small groups that like to share for the sake of sharing. They likely won’t care about the AI crawlers stealing their content, although I think there is a big overlap between people who still run blogs and those who dislike AI. 2. small organizations that are dedicated to one specific topic and are often largely ad financed. These organizations would likely stop to exist in such an AI search dominated world. > Especially since website hosting is close to being free these days. It is under specific circumstances. The problem is that those AI crawlers don’t check by once in a while like Google does but instead they hit the site very frequently. For a static site this won’t be much of an issue except for maybe bandwidth. For more complex sites like - say - the GitLab instances for OSS projects, reality paints a different picture
- birken 2y agoThe AI isn't "reading the web" though, they are reading the top hits on the search results, and are free-riding on the access that Google/Bing gets in order to provide actual user traffic to their sites. Many webmasters specifically opt their pages out of being in the search results (via robots.txt and/or "noindex" directives) when they believe the cost/benefit of the bot traffic isn't worth the user traffic they may get from being in the search results. One of my websites that gets a decent amount of traffic has pretty close to a 1-1 ratio of Googlebot accesses compared to real user traffic referred from Google. As a webmaster I'm happy with this and continue to allow Google to access the site. If ChatGPT is giving my website a ratio of 100 bot accesses (or more) compared to 1 actual user sent to my site, I very much should have to right to decline their access.
- jsbg 2y ago> If ChatGPT is giving my website a ratio of 100 bot accesses (or more) compared to 1 actual user sent to my site are you trying to collect ad revenue from the actual users? otherwise a chatbot reading your page because it found it by searching google and then relaying the info, with a link, to the user who asked for it seems reasonable
- birken 2y agoWhile yes, I am attempting to collect ad revenue from users, and yes, I don't want somebody competing with me and cutting me out the loop, a large part of it is controlling my content. I'm not arguing whether the AI chatbot has the legal right to access the page, I'm not a legal scholar. What I'm saying is that the leading search engines also have the equal rights to access whatever content they want, and yet they all give webmasters the following tools: - Ability to prevent their crawlers from accessing URLs via robots.txt - Ability to prevent a page from being indexed on the internet (noindex tag) - Ability to remove existing pages that you don't want indexed (webmaster tools) - Ability to remove an entire domain from the search engine (webmaster tools) It is really impolite for the AI chatbots to go around and flout all these existing conventions because they know that webmasters would restrict their access because it's much less beneficial than it is for existing search engines. In the long run, all this is going to lead to is more anti-bot countermeasures, more content behind logins (which can have legally binding anti-AI access restrictions) and less new original content. The victim will be all humans who aren't using a chatbot to slightly benefit the ones who are. And again, I'm not suggesting that AI chatbots should not be allowed to load webpages, just that webmasters should be able to opt out of it.
- GuinansEyebrows 2y agoSomeday I’ll have enough “karma” to downvote things like this. The agent should respect robots.txt no matter who is using the Robot.