5 ms·
This is 7 years old though, I would be curious to see if the experiment is still repeatable today. I've noticed personally that bing has gotten a lot better, b
by Someguywhatever 8y ago
This is 7 years old though, I would be curious to see if the experiment is still repeatable today.
I've noticed personally that bing has gotten a lot better, but that could just be because it's still a delayed version of Google.
Microsoft may even be using IE to harvest data to reverse engineer a Google search algorithm. Even if they make an inferior verision thats 90% as good, that's still good enough to keep them relevant in the market.
- exikyut 8y agoWell, you could sort of do this experiment yourself, with a bit of preparation and a decent bit of patience. - Prepare a new small website on an extremely obscure topic; something a search engine will consider uninteresting and not worth ranking highly, but just legitimate enough to warrant maintaining in the search index. On top of obscurity, don't link to the site from anywhere. This is so the site doesn't wind up crawled by Bing. - Manually add the site to Google's index (I forget where but there's a webmaster tool that can be used to manually add things; then you wait and see if the crawler paid any attention and added it to the index) - Add a single webpage on the site that, in amongst all the other words, contains a single random word, or a series of random words. Now repeatedly search Google for the term, on a machine nowhere near a Windows system, and and wait see if your site shows up. It will take some time to get this working, because unlike Google's fully contrived test, actually getting garbage/nonsense words into the index does take effort, as (in my experience) the index seems to be primed to index/prefer valid words over garbage; on top of this, your nonsense word will be (if you're doing it right) the only result in the world for the word, so it may be the search result is in the index but has some attribute marking it below a "show in results" threshold, if such metrics exist. - Now you have your magic term, do what Google did, and repeatedly search Bing for the term, from IE. - ??? - Profit?
- belltaco 8y agoThat experiment would have always failed. Check my other comment https://news.ycombinator.com/item?id=17810832 https://news.ycombinator.com/item?id=17810832
- exikyut 8y agoOnce the test site was successfully added to Google's index and the single webpage, and the nonsense term returned results for the test webpage, a toolbar-to-Google search from IE would just send the keyword off to Google, the search result would come back (from Google), and... Or I'm completely missing something. This is possible.
- belltaco 8y ago>Now you have your magic term, do what Google did, and repeatedly search Bing for the term, from IE That's not enough. You have to install the Bing toolbar on IE, and then consent to turn on the Suggested Sites feature.
- Someguywhatever 8y agoIn the actual experiment Google Engineers manually inserted the data into their index. IDK if my own activities would result in my fake site getting indexed by them, also I don't know how to ensure that I am the ONLY result, which was the case with the google test.
- exikyut 8y agoI've googled tons of things that only get one single search result. Obviously I can't link them here, since then there would be two results...
- CodingTom 8y agoI actually worked at Bing while all this was going down. It's a long time ago so I don't remember all the details, but when this surfaced it was a big deal internally. In fairness to Microsoft/Bing, the data collected as part of the Suggested Sites feature (and others in IE) was pretty clearly stated as being used to anonymously improve other Microsoft products, and so it became one of many sources of data for Bing's index. If you have a ton of data about user's traffic patterns on the web, you'd be crazy to not use it in an aggregate way to infer information about which URLs satisfy a user's intent at any given time. The irony is that nothing was even specifically targeted at using data from Google, the machine learning algorithms in question just naturally learned that people were typically satisfied by sites visited after searching for a term on Google, and therefore boosted those sites by some amount in Bing search results. No human was in the loop deciding that. In the contrived cases used by Google engineers in the blog post, the Google results were simply the ONLY sites being returned by the index for the given search term, as no other parts of the index found anything. After the issue was discovered, there was a huge push by the search relevance team to remove that data source. IIRC a new version of the index launched shortly afterwards that was stripped of all the user data features collected via IE, but maintained relevance parity through some amazing work by the team. I'd be extremely surprised if the experiment was repeatable today - I think they're much more careful about these things now. I don't think there's any intentional "reverse engineering" going on, the Bing team has a ton of engineers working on novel relevance techniques. At the time that I left, we were actually beating Google on our internal relevance metrics (and developing tougher new metrics in order to be able to keep measuring progress). EDIT: For evidence of this, one only needs to look at the volume of research that Microsoft publishes each year for SIGIR: https://www.microsoft.com/en-us/research/event/microsoft-sigir-2018/#!accepted-papers https://www.microsoft.com/en-us/research/event/microsoft-sig...