5 ms·
The thing that bothers me about all this drama is that the actual offense Google wants everyone to be so worked up about is that Bing doesn't filter Google from
by jfager 16y ago
The thing that bothers me about all this drama is that the actual offense Google wants everyone to be so worked up about is that Bing doesn't filter Google from its clickstream data.
Bing wrote code that works across the whole web. The whole web includes Google. As a result, Bing gets some info from Google. But they didn't get that info because they copied Google, they got it because they didn't filter Google out - or, said another way, because they ignored Google as someone they needed to special case for clickstream analysis.
I don't work in search, but the idea that you're supposed to special-case your competitors when writing general-purpose tools sounds an awful lot like a unilaterally recognized gentleman's agreement. If it's not illegal, and it doesn't hurt end users, why shouldn't it be considered fair game?
I also think it's odd that throughout this whole thing, nobody has really noticed that the only possible way Google could have spotted this issue is if they're keeping very close tabs on Bing's search results. It's another arbitrary line that Google seems to have unilaterally drawn: it's clearly fine to monitor your competitors results closely, which presumably is going to have an effect on your own results; it's only out of bounds when that effect is directly measurable.
- nostrademons 16y agoI thought part of the point is that whatever Bing is doing doesn't work across the whole web. They need to associate the URL with a query, and most websites don't have queries. It's not just that Bing has recorded a click on Miley Cyrus's webpage; it's that they've done that and associated it with the query [kecgxjpgqoe].
- jfager 16y agoPeople have made that point, but I don't understand it. 'kecgx...' shows up as a parameter in the url of the Google query. Tons of sites include relevant information in parameterized urls; why is it unexpected that Bing would use that information across the whole web? Other people have said that implies that Bing has to have special Google url-parsing code, but that's not true at all - query parameters in urls are standardized. You would have to have special code to understand the specific semantics of Google's query urls, but there's no reason to think Bing needs or wants parameter semantics, they could easily just be interested in making probabilistic associations.
- nostrademons 16y ago"You would have to have special code to understand the specific semantics of Google's query urls" That's why people say that it's special-cased. There's no web standard that says the 'q' parameter means that the page is a search engine and the parameter is the query. That was something AltaVista did a long time back (possibly for byte-saving reasons, or possibly because they were lazy) and Google et al copied. Many other search engines use a different system, eg. DDG puts it in the request path, InfoSeek used qt=, Excite used search=.
- jfager 16y agoWhy do you think Bing cares or has to know that "q=" means a Google query term? My point is that they don't have to have any semantic information about parameter keys to be able to derive probabilistic associations between parameter values and clicks. If you consistently see pages with 'foo' as a parameter value to any parameter key, and clicks on those pages consistently go to site bar, it's completely reasonable to start associating foo with bar, regardless of what the parameter keys are.
- Natsu 16y ago> Why do you think Bing cares or has to know that "q=" means a Google query term? If that's true, then they should also be associating the sites linked with all the other weird parameter values in a search query, which would spam them to heck. Here are all the params from a search I just did on google: q=test&ie=utf-8&oe=utf-8&aq=t&rls={moz:distributionID}:{moz:locale}:{moz:official}&client=firefox They'd start making a lot of strange associations between random sites and "utf-8" if that were true, because that parameter shows up in just about every Google search done in English. It's also a perfectly normal thing for programmers to search for, so they'd clutter up their index with millions of sites that had nothing whatsoever to do with utf-8. So to make any real use out of that, they had to understand what the parameters in there actually mean, rather than associating ALL of them with whatever site was next in the clickstream. Though I grant you, that does not disprove the alternate hypothesis that they were dumb and polluted their index with loads of irrelevant crap. And I admit that I found weak support for that hypothesis by trying to see if utf-8 was linked to rim.com (one of the tests sites, if memory serves): http://www.bing.com/search?q=utf-8+rim&go=&form=QBRE http://www.bing.com/search?q=utf-8+rim&go=&form=QBRE Those results appear to be crap, though I'm not sure that any sensible results exist that they could return.