3 ms·
This is a great idea. LLM crawlers are ignoring robots.txt, breaching site terms of service, and ingesting copyrighted data for training without a licence. We
by pyman 1y ago
This is a great idea.
LLM crawlers are ignoring robots.txt, breaching site terms of service, and ingesting copyrighted data for training without a licence.
We need more ideas like this!
- bhaney 1y agoThis is the same idea as in the article, just an alternative flavor of generating the zip bomb. And I actually only serve this to exploit scanners, not LLM crawlers. I've run a lot of websites for a long time, and I've never seen a legitimate LLM crawler ignore robots.txt. I've seen reports of that, but any time I've had a chance to look into it, it's been one of: - The site's robots.txt didn't actually say what the author thought they had made it say - The crawler had nothing to with the crawler it was claiming to be, it just hijacked a user agent to deflect blame It would be pretty weird, after all, for a company running a crawler to ignore robots.txt with hostile intent while also choosing to accurately ID itself to its victim.
- shakna 1y agoPerplexity certainly was ignoring robots.txt [0] Anthropic... Their robots.txt requires a delay to be defined, even though its an optional extension. But whatever. [0] https://www.wired.com/story/perplexity-is-a-bullshit-machine/ https://www.wired.com/story/perplexity-is-a-bullshit-machine...
- pyman 1y agoThere's plenty of evidence to the contrary; https://mjtsai.com/blog/2024/06/24/ai-companies-ignoring-robots-txt/ https://mjtsai.com/blog/2024/06/24/ai-companies-ignoring-rob...