3 ms·
I don't think it ignores robots.txt, I think it just doesn't have a very good parser and you need to give them their own user-agent block. I had a similar level
by tech2 3y ago
I don't think it ignores robots.txt, I think it just doesn't have a very good parser and you need to give them their own user-agent block. I had a similar level of frustration.
https://www.feitsui.com/en/article/32 https://www.feitsui.com/en/article/32
- codetrotter 3y agoAfter all, if they wanted to completely ignore the wishes of the website owners they probably would not announce their spider as such in the user agent. They’d just pretend to be a web browser.
- dotancohen 3y agoIt is trivial to detect a spider from human traffic based on requests alone. Lying about the UA would just be bad press for them.
- kccqzy 3y agoIf it's really trivial as you say, Google's reCAPTCHA and similar products like hCAPTCHA would instantly have no reason to exist.
- jpc0 3y agoBot intentionally trying to look human =! Spider A spider will generally have a pretty predictable route through a web site.
- dotancohen 3y agoThe various CAPTCHA implementations are primarily designed to prevent bot submissions, not spiders.
- codetrotter 3y agoSome of them yes. But not all. Try for example to browse a Cloudflare protected site from Tor and you will be hit with a constant barrage of captchas even though you are only doing GET requests.
- dotancohen 3y agoYes, huristicly, a tor browser is more likely to be nefarious than a regular browser user. Note the use of huristisc - such as IP address - not related to user agent.