7 ms·
Bing search results showing up in Google
- aristidb 16y agorobots.txt applies to the source of the links, not the target. So if, say, http://www.paulgraham.com/ http://www.paulgraham.com/ links to http://m.bing.com/search http://m.bing.com/search, then http://m.bing.com/robots.txt http://m.bing.com/robots.txt does not apply to that. EDIT: If you think this is wrong, please explain it instead of just downvoting me, because I think it is pretty unfair that I lose karma for explaining my interpretation.
- benologist 16y agoWhy would someone with no authority over your site linking to your site override your robots.txt? Edit: I didn't downvote you, but I think you're wrong because it makes no sense - if I link to your site that shouldn't give search engines a free pass to ignore your wishes and do whatever they want with your content.
- aristidb 16y agoI understand a robots.txt "Disallow: /foo" to mean that it must not crawl that page, i.e. look at the links _inside_ that page.
- benologist 16y agoI've always interpreted it to mean they're explicitly not allowed to touch it - no exceptions (unless you actually specified exceptions for them which you can see on the 3rd last example at http://www.robotstxt.org/robotstxt.html http://www.robotstxt.org/robotstxt.html).
- moultano 16y agoOne counter intuitive thing. They are allowed to link to it in search results (using links that point to it to rank it for insance) but can't use the content of the page.
- aristidb 16y agoInteresting, and I've always interpreted it the other way. Looks like you are right. But the language on www.robotstxt.org is pretty vague, too. (What does it mean to "visit" an URL?) The "RFC" on http://www.robotstxt.org/norobots-rfc.txt http://www.robotstxt.org/norobots-rfc.txt (Section 3.2.2) states: "These lines indicate whether accessing a URL that matches the corresponding path is allowed or disallowed. Note that these instructions apply to any HTTP method on a URL." And I think both interpretations are in theory (but probably not in practice) valid, at least if a search engine is willing to add a site to its index without accessing it (which is unlikely, but not impossible).
- Zakuzaa 16y agoThat's not true.
- pcowans 16y agoI think you're confusing robots.txt with the nofollow attribute value: http://en.wikipedia.org/wiki/Nofollow http://en.wikipedia.org/wiki/Nofollow
- gorset 16y agoyup. Anchor text is fair game. Keywords taken from anchor texts are usually a good indication on what the linked page is about. A page can have a high page rank without ever being crawled - given that other high ranked pages have linked to it with good keywords in the anchor text. Such pages can often be identified by the lack of snippet texts. Edit: I may have misunderstood you, you can only use the data found in links, not follow them.
- ajays 16y agoIf a page is not crawlable, but listed in the search results based on anchor text (or the URL), then it will not have any snippets. You can show snippets only if you crawl the page itself.
- ajays 16y agoIt just means a crawler cannot retrieve the page; in other words, a crawler cannot GET a URL for whom there's an exclusion pattern in robots.txt .
- jamesjyu 16y agoI agree with the author here. I think that Google will come away from this looking combative and childish.
- mwg66 16y agoBit different.
- moultano 16y agoThis looks like microsoft is assuming robots.txt is case insensitive?
- ajays 16y agoIt is. Only domain names are case-insensitive. This document explains how Google handles robots.txt : http://code.google.com/web/controlcrawlindex/docs/robots_txt.html http://code.google.com/web/controlcrawlindex/docs/robots_txt...
- deleted 16y ago[deleted]
- Natsu 16y agoAre we reading the same file? Under where it describes matching paths (so replace /fish and /Fish with /search and /Search if you like): ==Example path matches== [path] /fish Matches: /fish /fish.html /fish/salmon.html /fishheads /fishheads/yummy.html /fish.php?id=anything Does not match: /Fish.asp /catfish /?id=fish Comments: Note the case-sensitive matching.
- floatingatoll 16y agoDear Microsoft, case-sensitivity is important. [1] m.bing.com/robots.txt says "/search", not "/Search". All of the crawled [2] urls are "/Search" or "/~/search". Also, wap.bing.com/robots.txt explicitly "Allow:"s several search pages, which are indexed by google. [1] EDIT: Case insensitivity is often important. Above comment notes that some robots are case-insensitive. I suspect Google is not, based on the results. [2] EDIT: I said indexed, a reply corrects to crawled. Good point, thanks.
- marcog1 16y agoSmall correction: All crawled urls are "/Search" or "/~/search". Note that www.bing.com/search/ is indexed (link was found on another page), but not crawled (result has no snippet).
- amalcon 16y agoNaturally, Bing is hosted on a Windows server, which inherits the Windows filesystem eccentricities. Case insensitivity is among those. Because "search" has six letters, that would mean that the robots.txt would need to have 64 entries to completely exclude this directory. That's not even including the tilde thing or any other paths to that directory. And that's for one directory. Lame? Yes. Google's fault? Not in the slightest. But it brings up an interesting question: if MS clickstream gathering included an opt-out mechanism that happened to be impractical for Google, would that change the ethics of any of this? Say, by having the Bing toolbar identify itself in user-agent so that Google could block it if they wanted? I wouldn't think that would materially change the situation. If Google really wanted to, they could probably "block" this now by encrypting their existing URL redirects, thus hiding the URL from the Bing toolbar entirely, at least until the user is out of the Google system.
- xilun0 16y agoWhat is the value of processing robots.txt in a case-sensitive way? If urls are to have different status when considering case change, the the site structure is just broken. Plus considering robots.txt in case-sensitive way has already result in lots of errors, this one included, and will result in even more in the future. Plus HTTP is not mandated to use case-sensitive URL (though it's recommended). I can't think of any argument of why robots.txt should be processed in a case-sensitive way (i mean for a good reason -- obviously search engine have a very "good" incentive to handle it that way: the possibility to cheat and index more than they should, with an excuse when they are caught), on the other hand i can think of many for case-insensitive processing...
- deleted 16y ago[deleted]
- ashleyw 16y agoAren't /search and /Search considered two different directories when it comes to robots.txt?
- jfr 16y agoYes. RFC 3986, sections 6.2.2.2 and 6.2.3.
- nostrademons 16y agoI think you got the wrong RFC number: http://tools.ietf.org/rfc/rfc3938.txt http://tools.ietf.org/rfc/rfc3938.txt No mention of robots.txt, nor a section 6.
- jfr 16y agoSorry, the correct number was 3986. RFC 3986 - Uniform Resource Identifier (URI): Generic Syntax
- jedsmith 16y agoWith respect to the author, the conclusion here is very flawed. If you search for Bing in Google, you get Bing all over page 1. If you search for Google in Bing, you get Google all over page 1. That's not the result of Google capturing click stream data from Google Chrome and copying Bing's results, nor is it the result of Microsoft capturing click stream data from IE8 and copying Google's results. That's just the nature of indexing. As for robots.txt disallowing those URLs, there is no standard for robots.txt behavior. I have observed some user agents treat it as case insensitive, and others treat it as case sensitive. Honestly, this isn't even in the same ballpark as the Google accusations made earlier this week, and it smacks of just looking for things to accuse Google of in response to the "Binggate" (ugh, I typed it) drama. Can't we go back to more productive things?
- deleted 16y ago[deleted]
- unp3rs0n 16y agoI don't understand why everyone is using the term "copying the results". I think what Bing did was very smart, they incorporated user clickstream data. One could accuse this method of walking a thin line morally, but I suspect that Google's accusation wouldn't have stood any water as a lawsuit.
- luigi 16y agoBecause by incorporating clickstream data from Google, they're effectively copying Google search results. Bing should blacklist Google from its clickstream data.
- ajays 16y agoNot necessarily. It is possible that Google incorporates clickstream data too. The problem that Google's little experiment highlighted was: given the utter lack of any other signal, Bing uses the fact that that URL was ranked #1 by Google's search engine and clicked on by a user. Having said that: had I been doing this experiment at Google, I would have also added the following variations: - for some search terms, rank the honeypot URL #1 but don't click on it - for some search terms, rank the honeypot URL #1 on some _other_ search engine's list and click on it. How can they do that? There are search engines out there which use Google in the backend. Experiment #1 would have shown more blatant copying. Experiment #2 would have shown whether it's just Google, or any other search engine.
- Herring 16y ago>never mind that Bing only used its toolbar as a url discovery device, not to 'copy search results' Yeah, they just happened to discover high quality urls on google. What are the chances?
- ajays 16y agoURL discovery is one thing; ranking that discovered URL at number 1 without any other signals is another. I don't think anyone cares that much about how Bing does URL discovery (unless, of course, the URL is supposed to be private and exchanged via email). Given that they have a new URL, what made them rank it #1?
- mukyu 16y agoWe need to get over this partisan "gotcha journalism". No one really benefits from everyone making low content blog posts with any random accusations that make their side 'right' (which just happens to be ad hominem anyways).
- maeon3 16y agoMicrosoft is way out of line. Google figures out what content is good by crawling every page and doing the leg work, and Bing copies Google data and displays what google displays. Google proved it with the bing sting. there is absolutly NO reason why bing should have linked to those documents, other than that they copied off of Google's exam paper. When students do this, it is called plagiarizing. The smoke getting thrown by MS is just to distract and divert while they scramble to hide what they did.
- sid0 16y agoGoogle proved it with the bing sting They didn't prove anything. Every experiment needs a control. Where's Google's?
- ajays 16y ago"Control" what? Where's the control for "gravity pulled the apple down" ? It's called an existential proof. If you want to prove that a black swan exists, all you need is 1 example of a black swan.
- sid0 16y agoIt's called an existential proof. If you want to prove that a black swan exists, all you need is 1 example of a black swan. Strictly speaking, yes, non-controlled experiments exist. However without a control you cannot eliminate alternate explanations. Since we know that an alternate explanation exists here, the experiment doesn't show anything. In the case of the black swan there's no credible explanation for that thing you see being anything but a black swan. The control is essentially Occam's razor.
- Natsu 16y agoGoogle's hypothesis was "Does Bing use Google data in its rankings?" That hypothesis was proven (because there was no other way to get those crazy links except for clickstream data showing users going from a google search for those nonsense terms to those sites). If you want to explore a very different hypothesis, namely "Does Bing single out Google in its weightings of clickstream data?" then I suggest you go here: http://projectgus.com/files/googlebing/seaport-trace.txt http://projectgus.com/files/googlebing/seaport-trace.txt That's a packet capture of some clickstream data. That should be more than enough to forge as much data as you like. You can then make up a nonsense search term, like "doesb1ngtrustgmorethananyoneelse" that should get zero results, then forge clickstreams going from a google.com search to "yes.com" as well as an equal number of searches for that term going from some other search engine to "no.com" and you can explore the weights as much as you want. That said, I'd say that the fact that they weight Google highly enough that they'd take their word for it that a clearly irrelevant term should be mapped to some site is strong evidence, I think. Mind you, I don't think it's "wrong" exactly for Bing to do this. I'm not worried about it destroying search, either. The spammers/SEO types will make it useless soon enough.
- dminor 16y agoGoogle has likely indexed links to Bing found on other pages, rather than on Bing itself. That doesn't mean it followed the links (and it wouldn't, if excluded by robots.txt).
- pmb 16y agorobots.txt disallows (or did until recently) only "/search". The results shown have "/Search" in the url. Bing screwed up.
- haberman 16y agoNever underestimate the ability of a human being to rationalize. If I was in the Microsoft camp, I'm sure I would also be grasping at straws to explain why it's totally fine for Bing to use Google's search results. It's human nature to rationalize. The bottom line is that Bing's index contains associations that it could never have figured out if Google hadn't figured them out first. How many there are, we cannot know. There's no way around the fact that Bing is piggybacking on the work of Google's search engineers. Is it "good for the customer?" In the short term, it's good for the customer if they can buy $1 bootlegged DVDs. In the long term, it's bad for the consumer if the money goes to bootleggers instead of the people who are doing the actual work. Think I'm exaggerating the effect of just "1 out of 1000 signals?" This argument would be extremely easy to refute. Stop using Google's results. If it really isn't that significant, then why should it be a problem to stop using it? Just turn it off and let everyone observe that the quality is 99.9% as good as it used to be, and avoid any accusation of copying. By refusing to turn it off, Microsoft makes it clear that it is an important part of their index, and that they have no qualms about having an important part of their index ripped off wholesale from their biggest competitor. Maybe it's a smart business move. But if that's the case, spare us the outrage about being called "copyists."
- zaidf 16y agoBy that standard, Google should also quit trying to integrate invite-your-fb-friends feature, right? Cuz it's essentially facebook who has figured out how to get massive user signups with their real data and it's google trying to leech. Instead of not trying to integrate fb, Google is accusing fb of not opening up the data. So when it's convenient, you want data opened up. When it's not convenient, you scream "copycat!".
- jsnell 16y ago> never mind that Bing only used its toolbar as a url discovery device That is obviously untrue, and shows that the author does not understand the issue even superficially. The Google experiment showed that Bing was associating urls to search terms for no reason other than that Google had done so. You know, like making a search for mbzrxpgjys return rim.com, a URL which we can safely assume Bing was already quite aware of.
- ajays 16y agothe author does not understand the issue even superficially. That, in a nutshell, is it. I don't know why we're spending so much time on this post, as the author has no idea about how search engines work, and what is robots.txt . If he had just looked at Bing's robots.txt and the URLs in Google's results, he would have seen that each and every one of them passed robots.txt . Since he site-restricted the search to ".bing.com", naturally you _will_ get only Bing results!
- deleted 16y ago[deleted]
- illdave 16y agoIf you look at http://wap.bing.com/robots.txt http://wap.bing.com/robots.txt, the URLs that Google is returning are actually all set to 'allow', not disallow. It also looks like m.bing.com/robots.txt blocks /search while their actual URLs are /Search - I guess Googlebot treats robots.txt as case-sensitive.
- deleted 16y ago[deleted]
- shareme 16y agorobot.txt excludes /search not /Search..big difference as 99.99% of return results are ../Search* MS mistake on robot.txt file not Google's
- deleted 16y ago[deleted]
- Matt_Cutts 16y agoThe two major issues in this article were: - Google can see and return links to pages without crawling them. I made a video and a blog post about this a while ago: http://www.mattcutts.com/blog/robots-txt-remove-url/ http://www.mattcutts.com/blog/robots-txt-remove-url/ - URL paths are case sensitive. Bing blocks /search in its robots.txt but not /Search. That's how the /Search urls got crawled. In a later edit, the author suggests "It would be fairly trivial for bots to test if the server is IIS (if the server identifies itself as such of course) or to try to retrieve Robots.txt and robots.txt, if those come up as equal then the sever can be assumed to be case insensitive." The issue of case sensitivity in robots.txt is a long, very nuanced topic. Here's just one example to get you started: at least back in 2007 when we were talking about this amongst ourselves at Google, the web server for developer.apple.com was case-insensitive, but their robots.txt had lines like this: Disallow: /documentation/quicktime/ Disallow: /documentation/Quicktime/ Disallow: /documentation/QUICKTIME/ Disallow: /documentation/macosx/ Why would they do that? Apparently because Apple wanted the canonical link to be /documentation/QuickTime . Back then, at least 21M robots.txt files on the web had mixed-case paths. If Google started interpreting robots.txt files from servers that claimed to be IIS differently... well, I'll leave it as an exercise to the reader to come up with some of the unexpected bugs and behavior that could result. I know it's really tempting to write a headline like "Bing search results showing up in Google," but I wish the author had done more research instead of going for a gotcha. Any SEO worth his/her salt could have explained what was going on here.
- mc32 16y ago>"...but I wish the author had done more research instead of going for a gotcha." That's a bit of irony right there, isn't it? --wow, Matt, I guess it cuts both ways, eh?
- Matt_Cutts 16y agoAnything in particular you're talking about? I've said a lot of stuff online over the last 10 years. :)
- 16y ago
- yaix 16y agoShouldn't the headline be "Bing explicitly allowing some results pages to showing up in other search engines". The wap.bing.com/robots.txt blocks all /search/ and then explicitly allows a few. What ever the reason is for that. Very weak article, IMHO.