4 ms·
Is there a reason why you are not using robots.txt [1] to block it? [1]: https://developer.amazon.com/amazonbot#how-can-i-control-what-amazonbot-crawls-on-my-s
by kinduff 2y ago
Is there a reason why you are not using robots.txt [1] to block it?
[1]: https://developer.amazon.com/amazonbot#how-can-i-control-what-amazonbot-crawls-on-my-site https://developer.amazon.com/amazonbot#how-can-i-control-wha...
- dylan604 2y agoIs there a reason you feel that that file will be respected?
- lambdaxyzw 2y agoIs there a reason why you don't? Is it just general bitterness and cynicism? As far as I know all major search engines respect rebots.txt, I don't see why LLM scrappers would be different.
- dylan604 2y agoprobably, yes. But coming from LLM scrappers, I have absolutely no faith in any of them. When one of them calls themself "open" in their name and is anything but, why would I trust them for anything after they lie in their name? I also do not trust Google only crawls what is allowed in robots.txt. Maybe they only use the data allowed in public use, but I have no faith that they don't have crawled data in their version of shadow profiles. I do not trust bigTech at all, and for those that do, I really don't understand why you do.
- kinduff 2y agoBecause it's on their documentation. If OP had the file and entry, and they didn't respect it, then it would be another conversation.
- kemayo 2y agoBots from big companies like Amazon, which is who the author is complaining about, do tend to respect it. In fact, it's listed in their documentation that the GP linked to that they will. They could be lying -- but why bother?
- fragmede 2y agoAmazon's official documentation for Amazonbot, at https://developer.amazon.com/amazonbot https://developer.amazon.com/amazonbot states > Amazonbot respects standard robots.txt rules.