5 ms·
OpenAI, Anthropic ignoring rule that prevents bots scraping online content
- cdme 2y agohttps://archive.ph/bVgFO https://archive.ph/bVgFO
- kolinko 2y agoI wonder what’s the proof there
- dylan604 2y ago/var/log/http/access_log???
- MrThoughtful 2y agoIn the end this whole debate will come down to the question if a new law should be introduced. Something like "learnright", similar to todays "copyright". "learnright" would give the exclusive right to learn from a given piece of content to the author. So everybody else who wants to have their robots learn from it would need to make a deal with the author. My expectation is that this will not happen and robots will be allowed to learn from everything that is public. Just like humans.
- lagniappe 2y agoRespecting robots.txt isn't a rule, it's a courtesy.
- sschueller 2y agoIt can be a legal thing and hold up in court in some countries. Is respecting a "no trespassing" sign also just courtesy?
- dylan604 2y agoWhat countr[y|ies] has codified respecting robots.txt file in a legal manner?
- throwup238 2y ago> It can be a legal thing and hold up in court in some countries. Which countries? (I'm genuinely curious) > Is respecting a "no trespassing" sign also just courtesy? In countries with right to roam, the answer is often yes. In states like California with right to access waterways, many "no trespassing" signs are unenforceable too if they block access to rivers or beaches.
- Ukv 2y ago> Which countries? (I'm genuinely curious) The EU's AI act points to the DSM directive's text and data mining exemption, allowing for commercial data mining so long as machine-readable opt-outs are respected - robots.txt is typically taken as the established standard for this. It has generally been respected as far as I'm aware (though, this article asserts otherwise). In the US, respecting robots.txt/<NoAI>/... could potentially be a fallback if the Fair Use defense falls through, by implied license like for website caching in Field v. Google Inc ("Google reasonably interpreted absence of meta-tags as permission to present 'Cached' links to the pages of Field's site").
- deleted 2y ago[deleted]
- deleted 2y ago[deleted]
- tedivm 2y agoIn the US, where all these companies are based, it is legal to scrape any site that doesn't require a login regardless of the robots.txt file. However, what you do with that data is still subject to copyright law.
- OutOfHere 2y agoIt is akin to a religious convention that some believe in. If you force your religion on others, expect others to force theirs on you, so don't. It is not trespassing because the site is already on the public web.
- skybrian 2y agoCourtesy is a practical way of avoiding legal disputes. If you ignore robots.txt, that means you don’t have an agreement, so the question is what do you have a right to do, and do you really want to argue about that?
- lagniappe 2y agoYou're free to bring anything to court, that's your right.
- alwa 2y agoI’m suspicious over the scant detail in this claim. I feel like when this has come up in the past, it’s turned out that the training scrapes obeyed robots.txt, but the agents browsing in direct response to user conversations were retrieving page content to summarize, in the spirit of following the link on a search result page. Maybe still something people would prefer not happen, but very much ephemeral access in response to an individual user’s query, not training models. There’s nothing in the linked article to distinguish between those types of requests or usage, and no evidence to suggest the firms are doing other than what they say they are.
- ipnon 2y agoIt’s not an HTTP demand, it’s an HTTP request.
- Ukv 2y ago> The world's top two AI startups are ignoring requests by media publishers to stop scraping their web content for free model training data, Business Insider has learned. Especially with the rise of SEO/GenAI spam content, I kind of wish media had more rigor on citing claims - rather than leaving them as bare assertions - ideally so that someone could follow the chain all the way back and see the actual evidence giving rise to (and hopefully substantiating) the claim. Is this from first-hand investigation where you could show access logs? Is this further communication put out by "TollBit", the licensing broker, in addition to their mentioned accusations that didn't include names? Is this coming from anonymous insiders at OpenAI/Anthropic who emailed BI? Did this information materialise in front of the article's author?
- Xenoamorphous 2y agoIMO robots.txt rules should be binary, either all are allowed or none. None of that “only Google” nonsense.
- tsujamin 2y agoGenuinely curious as to why? If it’s your work, information or copyright, should you not have a say in how it is used? If you want your site to be indexed and discoverable by your audience, but not used as a source in some LLM summary, is that not your prerogative? Unless you mean in a more political manner, that copyleft should be prescribed to website owners
- JimDabell 2y ago> If it’s your work, information or copyright, should you not have a say in how it is used? In general, no. Copyright is the exclusive right to copy. If you buy a copy of a bestselling book and make copies of it, the copyright holder has the right to sue you, because the government has awarded them an exclusive (limited, temporary) monopoly over creating new copies of that work. If you buy a copy of a bestselling book and learn things from it, deface it, make origami out of it, burn it to keep warm over winter, or whatever, the copyright holder doesn’t get any say in the matter, because it’s copyright not useright.
- JimDabell 2y agoIf you aren’t recursively fetching pages, you aren’t a robot and shouldn’t follow robots.txt: > WWW Robots (also called wanderers or spiders) are programs that traverse many pages in the World Wide Web by recursively retrieving linked pages. — http://www.robotstxt.org/orig.html http://www.robotstxt.org/orig.html When these companies crawl the web looking for training data, they should obey.robots.txt. When a user asks them about a particular page (e.g. “summarise this page”), they should ignore robots.txt All of these articles purport to show that these companies are not doing the former by testing the behaviour of the latter. That’s a misunderstanding of how robots.txt is supposed to work.
- rolph 2y agoon spider request send robots ToS clickthrough on fail redirect.
- cdme 2y agoI just 403 them as best I can.