4 ms·
Can’t you just ignore robots.txt?
by adamhearn 6y ago
Can’t you just ignore robots.txt?
- saimiam 6y agoWell behaved crawlers don't. Products like Google Search and Bing have a self-interest in supporting robots.txt because the alternative is annoyed webmasters suing them for not providing an opt-out due to which their unsecured credentials are a quick search away. In fact, I believe the robots.txt protocol was devised in consultation with search engines.
- mrkramer 6y agoA lot of webmasters are amateurs who expose their credentials even when robot rules are respected. Google dorking is a way to find that credentials and other sensitive data and information.
- mxxx 6y agoUnless I’m mistaken, google don’t actually respect robots.txt any more. They recommend the use of some meta tag instead, from memory.
- yarcob 6y agoTheir docs say they do: https://developers.google.com/search/reference/robots_txt https://developers.google.com/search/reference/robots_txt They do say that they ignore robots.txt for user-initiated actions eg. Google Translate, which makes sense in my opinion.
- mxxx 6y agoAh yes you’re right, my apologies. They don’t support the noindex directive inside a robots.txt any more.