4 ms·
I don't know much about the content of robots.txt, can you share some mistakes they are making here?
by ewidar 5y ago
I don't know much about the content of robots.txt, can you share some mistakes they are making here?
- dazc 5y agoIt's not really a good idea to restrict crawlers at directory level since it doesn't prevent any of those pages being indexed. You can get weird behaviour such as googlebot trying to crawl a page but not being allowed to. If this happens a lot then they are going to penalise you for it because they don't want a bunch of pages in their index that they don't know the content of. If this number of pages is a significant percentage of your site then you are in real trouble. Much better to use x-robots noindex, nofollow for pages you don't want to be in the public domain. https://developers.google.com/search/docs/advanced/crawling/block-indexing https://developers.google.com/search/docs/advanced/crawling/...
- zinekeller 5y ago... if you're specifically targeting Google. If you are a neutral site (in this case, caters to a lot of people - especially in Asia, where Google as a search engine is a hit-or-miss proposition), then this might backfire sine other crawlers won't understand X-Robots. You can mitigate this by knowing which spiders knows X-Robots (Google, Bing and Yandex AFAIK) and whitelist them while still using a disallow directive for the rest of them.