5 ms·
Hmm, I'm not sure you RTFA. He didn't search for the term "bing", he searched for the term "site:.bing.com/search". So it IS kinda a big deal that they're disre
by anonymous246 16y ago
Hmm, I'm not sure you RTFA. He didn't search for the term "bing", he searched for the term "site:.bing.com/search". So it IS kinda a big deal that they're disrespecting the robots.txt and listing those pages anyway.
I have also seen Google listing one of my domains for which I specifically disallowed all spiders (Bing doesn't show those domains FYI). My feeble attempt at separating my personal and professional personas was defeated by frigging Google's blatant disregard for Internet etiquette.
Since I'm trying to be anon here, I won't be able to list the search term. Sorry.
- jedsmith 16y ago> So it IS kinda a big deal that they're disrespecting the robots.txt and listing those pages anyway. Did you RMFC? The robots.txt doesn't match what is indexed. As has been exhaustively pointed out in this thread by more than a few commenters, "Search" != "search". On that very search results page, we see #1 which is: > OLAC search - Bing OLAC is an unrelated site, and at one point it apparently linked to "www.bing.com/search" with the anchor text "OLAC search - Bing", which Google faithfully indexed but absolutely did not fetch, as you notice there is no description. What else would "OLAC search" mean? This gives us insight into how Google indexes and that they actually respect robots.txt case-sensitively. The second URL, m.bing.com/Search/Results.aspx, was indexed and fetched because "Search" != "search".
- xilun0 16y ago> "Search" != "search" This is highly ridiculous. HTTP URLs are not mandated to be case-sensitive (though it's recommended), and clearly lots of sites use them in a case-insensitive way. Robots should either consider robots.txt in a case-insensitive way (even if I'm conscious that lot of them, including major ones, currently don't do that, which is precisely what I consider to be a problem, and this is supported by what happened here -- where google is risking to appear as a fool). The following article has perfectly good arguments in favor of case-insensitivity or even more clever handling: http://www.slicksurface.com/blog/2007-04/be-careful-robotstxt-is-case-sensitive http://www.slicksurface.com/blog/2007-04/be-careful-robotstx... Also; failing to respect conditions of use of a service because an automatic process is not safe enough is not a completely exonerating excuse for the operator of such process...
- deleted 16y ago[deleted]
- prodigal_erik 16y agoAn HTTP client must always assume URL paths and query strings are case-sensitive. You can't rely on the bing.com web servers always returning the same resource for http://www.bing.com/search?q=test http://www.bing.com/search?q=test and http://www.bing.com/Search?q=test http://www.bing.com/Search?q=test just because the default filesystem for their OS is case-insensitive today. Actually it's unlikely the main entry point for their search engine is a file named "search" in some top-level directory. The only way for an author to tell you many URLs are equivalent is to use "301 Moved Permanently" responses to the canonical one from all the others, and Bing doesn't do that.
- haberman 16y agoThey are not disrespecting robots.txt: http://news.ycombinator.com/item?id=2183519 http://news.ycombinator.com/item?id=2183519 > I have also seen Google listing one of my domains for which I specifically disallowed all spiders [...] Since I'm trying to be anon here, I won't be able to list the search term. Sorry. Unsubstantiated accusations. If you truly think Google is disregarding robots.txt but you don't want to divulge the original site, you should set up an experiment as Google has. Otherwise we have no way of determining if your accusations are true, and therefore they should not be taken seriously.
- anonymous246 16y agoFine. I reason I stated my point was so that others could chime in if they've had a similar experience. Let me ask you and others this: Is the following robots.txt supposed to exclude all pages from my domain from showing up in Google results? Am I missing something? According to http://www.robotstxt.org/robotstxt.html http://www.robotstxt.org/robotstxt.html I think I'm doing the right thing. Same file is returned for www.<domain>.com/robots.txt and <domain>.com/robots.txt. Google lists <domain>.com/<subdir> in results. I don't think it should be. User-agent: * Disallow: /
- haberman 16y ago> Let me ask you and others this: Is the following robots.txt supposed to exclude all pages from my domain from showing up in Google results? I believe that robots.txt is a way to prevent your site from being crawled by a robot, but it is not a blacklist against your site appearing in Google search results if it finds a link to your page on a site that does allow robots. Check out this page: http://www.google.com/support/webmasters/bin/answer.py?hl=en&answer=164734 http://www.google.com/support/webmasters/bin/answer.py?hl=en... Specifically check out the section "I want to completely remove a page from search results." It appears that if you use the "noindex" meta tag, you can prevent the site from showing up in search results even if other pages link to it. The noindex meta tag is documented here: http://www.google.com/support/webmasters/bin/answer.py?answer=79812 http://www.google.com/support/webmasters/bin/answer.py?answe...
- illdave 16y agoThere's a bit of a weird myth with robots.txt and the idea that it prevents pages from showing up in search engine indexes. Robots.txt means that the search engine cannot crawl the page - it can still include it in the search results if it sees enough sites linking to it. It can take a guess at what the page title can be, but there's usually no descriptive snippet because it's unable to see what's on the page. If you don't want the page to be included in the index at all, you can use the meta noindex tag, by putting this in the head: <meta name="robots" content="noindex" /> Pro tip: you need to also unblock that page from robots.txt - if Google isn't allowed to crawl the page, it can't see the meta noindex tag, which means it would stay indexed.
- anonymous246 16y agoThanks. I think this may be the reason. IMHO, Bing is being more reasonable here. So, I have to let Google see any page which I want to prevent it from showing it to others. If I want Google to show NONE of my pages to others, I should show them ALL my pages. Conveniently, there is no wildcarded noindex, is there? Nonsensical, but since Google has more power here I'll probably have to bend to their whim.