5 ms·
People have made that point, but I don't understand it. 'kecgx...' shows up as a parameter in the url of the Google query. Tons of sites include relevant infor
by jfager 16y ago
People have made that point, but I don't understand it. 'kecgx...' shows up as a parameter in the url of the Google query. Tons of sites include relevant information in parameterized urls; why is it unexpected that Bing would use that information across the whole web? Other people have said that implies that Bing has to have special Google url-parsing code, but that's not true at all - query parameters in urls are standardized. You would have to have special code to understand the specific semantics of Google's query urls, but there's no reason to think Bing needs or wants parameter semantics, they could easily just be interested in making probabilistic associations.
- nostrademons 16y ago"You would have to have special code to understand the specific semantics of Google's query urls" That's why people say that it's special-cased. There's no web standard that says the 'q' parameter means that the page is a search engine and the parameter is the query. That was something AltaVista did a long time back (possibly for byte-saving reasons, or possibly because they were lazy) and Google et al copied. Many other search engines use a different system, eg. DDG puts it in the request path, InfoSeek used qt=, Excite used search=.
- jfager 16y agoWhy do you think Bing cares or has to know that "q=" means a Google query term? My point is that they don't have to have any semantic information about parameter keys to be able to derive probabilistic associations between parameter values and clicks. If you consistently see pages with 'foo' as a parameter value to any parameter key, and clicks on those pages consistently go to site bar, it's completely reasonable to start associating foo with bar, regardless of what the parameter keys are.
- Natsu 16y ago> Why do you think Bing cares or has to know that "q=" means a Google query term? If that's true, then they should also be associating the sites linked with all the other weird parameter values in a search query, which would spam them to heck. Here are all the params from a search I just did on google: q=test&ie=utf-8&oe=utf-8&aq=t&rls={moz:distributionID}:{moz:locale}:{moz:official}&client=firefox They'd start making a lot of strange associations between random sites and "utf-8" if that were true, because that parameter shows up in just about every Google search done in English. It's also a perfectly normal thing for programmers to search for, so they'd clutter up their index with millions of sites that had nothing whatsoever to do with utf-8. So to make any real use out of that, they had to understand what the parameters in there actually mean, rather than associating ALL of them with whatever site was next in the clickstream. Though I grant you, that does not disprove the alternate hypothesis that they were dumb and polluted their index with loads of irrelevant crap. And I admit that I found weak support for that hypothesis by trying to see if utf-8 was linked to rim.com (one of the tests sites, if memory serves): http://www.bing.com/search?q=utf-8+rim&go=&form=QBRE http://www.bing.com/search?q=utf-8+rim&go=&form=QBRE Those results appear to be crap, though I'm not sure that any sensible results exist that they could return.
- jfager 16y agoIf that's true, then they should also be associating the sites linked with all the other weird parameter values in a search query If it's a parameter value people will ever actually use Bing to search for ('utf-8'), there's probably plenty of other signal to help them figure out which results to return. If it's not ('kecgxjpgqoe'), we already know they sometimes return crap, thanks to Google's little experiment. They'd start making a lot of strange associations between random sites and "utf-8" if that were true, because that parameter shows up in just about every Google search done in English. If it shows up as a parameter for every search, why do you think Bing's algorithm would decide it was a good indicator for a particular url? P(foo.com|utf-8) wouldn't be any different from P(bar.com|utf-8), making 'utf-8' basically worthless as a discriminator. I'm pretty sure the folks at Bing understand the concept of conditional probability. It's also a perfectly normal thing for programmers to search for, so they'd clutter up their index with millions of sites that had nothing whatsoever to do with utf-8 I see no justification for the idea that url parameters are somehow a more difficult challenge in this regard than the mass of crap that is content on the web, which we know they crawl and index at a massive scale.
- Natsu 16y agoIt would only show up as a parameter "relevant" to whatever sites were next in the clickstream. Remember, not every page transition is Google -> other site. They'd also be gobbling up tons of random things from forums and whatnot (which should appear on long tail searches, if we knew where to look), most of which spam the heck out of you with random parameters, forum names, and whatnot. > I see no justification for the idea that url parameters are somehow a more difficult challenge in this regard than the mass of crap that is content on the web, which we know they crawl and index at a massive scale. Which is why I see no reason to assume that they don't understand or attempt to understand the actual meaning of parameters passed to one of the biggest sites on the internet.
- jfager 16y agoYou said 'utf-8' shows up as a parameter in almost every English Google search, and suggested that this would cause weird associations between 'utf-8' and random pages. I pointed out that there's no reason to expect this to be true, because Bing engineers are likely smart enough to realize that if the probability of clicking on to foo.com is not significantly different than the probability of clicking on to bar.com given the presence of the 'utf-8' parameter, then 'utf-8' is a pretty poor discriminator between foo.com and bar.com, and probably shouldn't be used to determine search results. It doesn't matter that not every page transition is Google -> another site. You still wouldn't need to special-case Google to determine an association with a parameter is useful or useless - the same code could build that model for any site with params. They'd also be gobbling up tons of random things from forums and whatnot (which should appear on long tail searches, if we knew where to look), most of which spam the heck out of you with random parameters, forum names, and whatnot. How do you know that they don't? Google pointed out some longtail results that look bad, and you yourself pointed some out in a previous comment. Which is why I see no reason to assume that they don't understand or attempt to understand the actual meaning of parameters passed to one of the biggest sites on the internet. You're misunderstanding what I'm saying. I don't assume that they don't; I'm just saying there's no evidence that they do, and that assertions of wrongdoing based on the belief that they do are just irresponsible speculation. I have no knowledge of what Bing actually does, but neither do the vast majority of the people on the internet who are talking about this, many of whom assume the worst based on a mistaken notion of what's technically necessary to see the results that Google demonstrated.