3 ms·
This is the same idea as in the article, just an alternative flavor of generating the zip bomb. And I actually only serve this to exploit scanners, not LLM cra
by bhaney 1y ago
This is the same idea as in the article, just an alternative flavor of generating the zip bomb.
And I actually only serve this to exploit scanners, not LLM crawlers.
I've run a lot of websites for a long time, and I've never seen a legitimate LLM crawler ignore robots.txt. I've seen reports of that, but any time I've had a chance to look into it, it's been one of:
- The site's robots.txt didn't actually say what the author thought they had made it say
- The crawler had nothing to with the crawler it was claiming to be, it just hijacked a user agent to deflect blame
It would be pretty weird, after all, for a company running a crawler to ignore robots.txt with hostile intent while also choosing to accurately ID itself to its victim.
- shakna 1y agoPerplexity certainly was ignoring robots.txt [0] Anthropic... Their robots.txt requires a delay to be defined, even though its an optional extension. But whatever. [0] https://www.wired.com/story/perplexity-is-a-bullshit-machine/ https://www.wired.com/story/perplexity-is-a-bullshit-machine...
- pyman 1y agoThere's plenty of evidence to the contrary; https://mjtsai.com/blog/2024/06/24/ai-companies-ignoring-robots-txt/ https://mjtsai.com/blog/2024/06/24/ai-companies-ignoring-rob...