4 ms·
I actually worked at Bing while all this was going down. It's a long time ago so I don't remember all the details, but when this surfaced it was a big deal inte
by CodingTom 8y ago
I actually worked at Bing while all this was going down. It's a long time ago so I don't remember all the details, but when this surfaced it was a big deal internally.
In fairness to Microsoft/Bing, the data collected as part of the Suggested Sites feature (and others in IE) was pretty clearly stated as being used to anonymously improve other Microsoft products, and so it became one of many sources of data for Bing's index. If you have a ton of data about user's traffic patterns on the web, you'd be crazy to not use it in an aggregate way to infer information about which URLs satisfy a user's intent at any given time. The irony is that nothing was even specifically targeted at using data from Google, the machine learning algorithms in question just naturally learned that people were typically satisfied by sites visited after searching for a term on Google, and therefore boosted those sites by some amount in Bing search results. No human was in the loop deciding that. In the contrived cases used by Google engineers in the blog post, the Google results were simply the ONLY sites being returned by the index for the given search term, as no other parts of the index found anything.
After the issue was discovered, there was a huge push by the search relevance team to remove that data source. IIRC a new version of the index launched shortly afterwards that was stripped of all the user data features collected via IE, but maintained relevance parity through some amazing work by the team. I'd be extremely surprised if the experiment was repeatable today - I think they're much more careful about these things now.
I don't think there's any intentional "reverse engineering" going on, the Bing team has a ton of engineers working on novel relevance techniques. At the time that I left, we were actually beating Google on our internal relevance metrics (and developing tougher new metrics in order to be able to keep measuring progress). EDIT: For evidence of this, one only needs to look at the volume of research that Microsoft publishes each year for SIGIR: https://www.microsoft.com/en-us/research/event/microsoft-sigir-2018/#!accepted-papers https://www.microsoft.com/en-us/research/event/microsoft-sig...