4 ms·
Fine. I reason I stated my point was so that others could chime in if they've had a similar experience. Let me ask you and others this: Is the following robots
by anonymous246 16y ago
Fine. I reason I stated my point was so that others could chime in if they've had a similar experience.
Let me ask you and others this: Is the following robots.txt supposed to exclude all pages from my domain from showing up in Google results? Am I missing something? According to http://www.robotstxt.org/robotstxt.html http://www.robotstxt.org/robotstxt.html I think I'm doing the right thing. Same file is returned for www.<domain>.com/robots.txt and <domain>.com/robots.txt. Google lists <domain>.com/<subdir> in results. I don't think it should be.
User-agent: *
Disallow: /
- haberman 16y ago> Let me ask you and others this: Is the following robots.txt supposed to exclude all pages from my domain from showing up in Google results? I believe that robots.txt is a way to prevent your site from being crawled by a robot, but it is not a blacklist against your site appearing in Google search results if it finds a link to your page on a site that does allow robots. Check out this page: http://www.google.com/support/webmasters/bin/answer.py?hl=en&answer=164734 http://www.google.com/support/webmasters/bin/answer.py?hl=en... Specifically check out the section "I want to completely remove a page from search results." It appears that if you use the "noindex" meta tag, you can prevent the site from showing up in search results even if other pages link to it. The noindex meta tag is documented here: http://www.google.com/support/webmasters/bin/answer.py?answer=79812 http://www.google.com/support/webmasters/bin/answer.py?answe...
- gojomo 16y agoSo it appears the only way to not appear is to let them crawl to see your NOINDEX tag. And while it's clear NOINDEX prevents a page from appearing in results, it's not clear that it excludes the page contents from analysis by any of Google's algorithms, once collected. (Is it still used to train the spell-checker, for example?)
- anonymous246 16y agoGreat. This link helps. Google's site basically says that in addition to specifically excluding their crawler via robots.txt, I also have MANUALLY submit a request to them. As I said in a reply below, this is nonsensical and Bing is being more reasonable, but I'll grit my teeth and do it since Google has more power here. It definitely seems like Google is exploiting a loophole in spirit of the definition of robots.txt. Robots.txt is an ancient standard, and I don't think it was anticipated at that time that search engines would gain enough confidence about pages' relevance to list them even if they had not indexed/crawled them.
- prodigal_erik 16y agoAs I recall, the spirit of robots.txt was not about appropriateness of search results so much as "this URL space can generate an unbounded graph, please don't DoS my server by trying to exhaustively traverse it."
- Matt_Cutts 16y agoNo. It prevents pages from being crawled, but references to uncrawled pages can still be shown. See http://www.mattcutts.com/blog/robots-txt-remove-url/ http://www.mattcutts.com/blog/robots-txt-remove-url/ for how it works.