3 ms·
I do not have a disallow set for this page but based off of this page it seems perhaps I might have to set one: https://www.deepcrawl.com/blog/best-practice/no
by scalesolved 8y ago
I do not have a disallow set for this page but based off of this page it seems perhaps I might have to set one:
https://www.deepcrawl.com/blog/best-practice/noindex-disallow-nofollow/ https://www.deepcrawl.com/blog/best-practice/noindex-disallo...
Specifically the part:
>Noindex (robots.txt) + Disallow: This prevents pages appearing in the index, and also prevents the pages being crawled. However, remember that no PageRank can pass through this page.
- detaro 8y agoGoogle specifically warns that having a Noindex header doesn't work if there is robots.txt disallow (since the crawler never sees the noindex, since it obeys the disallow), that's why I asked.
- r1ch 8y agoThere's a non standard Noindex: /URL that Google accepts in robots.txt. Unfortunately it breaks many other crawlers which don't understand it, so you have to rely on user agent sniffing.