3 ms·
I wrote about this a while ago [1]. Unfortunately, with robots.txt, you're at the mercy of crawlers. They may respect it or ignore it altogether. You can block
by microflash 3y ago
I wrote about this a while ago [1]. Unfortunately, with robots.txt, you're at the mercy of crawlers. They may respect it or ignore it altogether. You can block IP addresses but many crawlers may not even use static IP addresses.
You can go to extremes and put your content behind a login, as others have suggested. But that would also create friction for your intended audience.
[1]: https://www.naiyerasif.com/post/2023/09/30/blocking-ai-web-crawlers/ https://www.naiyerasif.com/post/2023/09/30/blocking-ai-web-c...
- Slavaqua 3y agoIt sounds like a loosing battle to manually keep track of all bots and their deployment IPs that generate datasets that LLM's might use or start using in the future. There must be a legal solution, a licence that forbids the use of content for training without explicit permission from the author.