5 ms·
What's wrong with AI agents accessing website content? We seem to have been happy with Google doing that for ages in exchange for displaying the website in sea
by zkid18 2y ago
What's wrong with AI agents accessing website content?
We seem to have been happy with Google doing that for ages in exchange for displaying the website in search results.
- spiderfarmer 2y agoAnd AI agents scrape your content in exchange for what exactly?
- zkid18 2y agoSorry, I distinguish here an AI agent that basically automate the visual lookup and scraping to feed into LLMs by big tech. I don't see any problem with the first one tbh.
- deleted 2y ago[deleted]
- red_admiral 2y agoThe website owner chooses. They can say "nope" in robots.txt. Not everyone respects this, but Google does. Google can choose not to show that site as a result, if they want to. This adds a third option besides yes and no, which is "here's my price". Also, because cloudflare is involved, bots that just ignore a "nope" might find their lives a bit harder.
- lolinder 2y agoRobots.txt is for crawlers. It's explicitly not meant to say one-off requests from user agents can't access the site, because that would break the open web.
- Spivak 2y agoYep, there's really two parts to this. * Some company's crawler they're planning to use for AI training data. * User agents that make web requests on behalf of a person. Blocking the second one because the user's preferred browser is ChatGPT isn't really in keeping with the hacker spirit. The client shouldn't matter, I would hope that the web is made to be consumed by more than just Chrome.
- lolinder 2y agoYeah, there's a lot of confusion between AI training and AI agent access, and it's dangerous. Training embeds the data into the model and has copyright implications that aren't yet fully resolved. But an AI agent using a website to do something for a user is not substantially different than any other application doing the same. Why does it matter to you, the company, if I use a local LLaMA to process your website vs an algorithm I wrote by hand? And if there is no difference, are we really comfortable saying that website owners get a say in what kinds of algorithms a user can run to preprocess their content?
- jsheard 2y ago> But an AI agent using a website to do something for a user is not substantially different than any other application doing the same. If the website is ad-supported then it is substantially different - one produces ad impressions and the other doesn't. Adblocking isn't unique to AI agents of course but I can see why site owners wouldn't want to normalize a new means of accessing their content which will inherently never give them any revenue in return.
- lolinder 2y agoI don't believe that companies have the right to say that my user agent must run their ads. They can politely request that it does and I can tell my agent whether to show them or not.
- jsheard 2y agoTrue, but by the same measure your user agent can politely request a webpage and the server has the right to say 403 Forbidden. Nobody is required to play by the other parties rules here.
- lolinder 2y agoExactly. The trouble is that companies want the benefits of being on the open web without the trade-offs. They're more than welcome to turn me down entirely, but they don't do that because that would have undesirable knock-on effects. So instead they try to make it sound like I have a moral obligation to render their ads.
- brigadier132 2y agoFor traditional search indexing the interests of the aggregator and the content creator were aligned. AIs on the other hand are adversarial to the interest of content creators, a sufficiently advanced AI can replace the creator of the content it was trained on.
- lolinder 2y agoWe're talking in this subthread about an AI agent accessing content, not training a model on content. Training has copyright implications that are working their way through courts. AI agent access cannot be banned without fundamentally breaking the User Agent model of the web.
- brigadier132 2y agoOk, fine, let's restrict it to AI agents only, without training. It's still an adversarial relationship with the content creator. When you take an AI agent an ask it "find me the best italian restaurant in city xyz" it scans all the restaurant review sites and gives you back a recommendation. The content creator bears all the burden of creating and hosting the content and reaps non of the reward as the AI agent has now inserted itself as a middleman. The above is also a much clearer / more obvious case of copyright infringement than AI training. > AI agent access cannot be banned without fundamentally breaking the User Agent model of the web. This is a non-sequitur but yes you are right, everything in the future will be behind a login screen and search engines will die.
- lolinder 2y ago> reaps non of the reward Just to be clear what we're talking about: the reward in question is advertising dollars earned by manipulating people's attention for profit, right? I frankly don't think that people have the right to that as a business model and would be more than happy to see AI agents kill off that kind of "free" content.
- brigadier132 2y ago
- 6gvONxR4sf7o 2y agoThe thing people have been doing for ages is a trade: I let you scrape me and in return you send me relevant traffic. The new choice isn't about a trade, so it's different.