21 ms·
"Copying" was a pretty brutal word to use-- not surprising that it raised MSFT's hackles a bit. MS clearly uses toolbar users' clickstreams (on and off Google)
by webwright 16y ago
"Copying" was a pretty brutal word to use-- not surprising that it raised MSFT's hackles a bit.
MS clearly uses toolbar users' clickstreams (on and off Google) to improve their own search efforts. Google created an artificial scenario where the ONLY input was Google search behavior and lo, the search results are exactly the same. Whether or not that steps over a line (I don't feel that it does), it's not "copying" in my book.
Another interesting point is that Google has been beating the "open" drum for a long time now. No walled gardens, right? If a Facebook user should have the right to take his data with him wherever he wants to go, shouldn't a Google user be able to fork over their behavior data to Bing?
Matt's point about MS' lack of clarity when getting folks permission to grab their clickstream was dead-on. THAT is pretty outrageous and MS should be ashamed of that.
Regardless of all that, hats off to Matt for keeping a cool head and stating his position in a respectful way.
- jimbokun 16y ago"Google created an artificial scenario where the ONLY input was Google search behavior and lo, the search results are exactly the same." Cutts addresses this: "As we said in our blog post, the whole reason we ran this test was because we thought this practice was happening for lots and lots of different queries, not simply rare queries."
- contextfree 16y agoIf Google think they have other evidence, let's see it. The burden of proof is on the accuser, I think. (c/p from other thread.)
- rayval 16y agoThe problem is that the results would not be so dramatic. It would require statistical analysis to show that there was a difference. Most people's eyes would glaze over -- even though copying was occurring and genuine damage was being done (to market share, to business model, etc.), similar to a low-grade toxin in the environment.
- jerf 16y agoIf this isn't enough to convince you, nothing will be. I'm not sure how the evidence could be any more convincing. Content from Google's result database ends up used to generate Bing's results, full stop. Proved. It may be technical error, it may be only used a little bit, it may be deliberate, it may be accidental and perfectly ethical by some people's standards, there's all kinds of ways to legitimately interpret this fact. But Google has demonstrated beyond a shadow of a doubt that input from Google is used to generate some non-zero bit of Bing's results. There is absolutely, positively no other explanation that accounts for the facts of the case.
- contextfree 16y ago"... we thought this practice was happening for lots and lots of different queries, not simply rare queries ..." My comment was in reference to that.
- jerf 16y agoI know. My point is that they've established beyond a shadow of a doubt that Google information is ending up in Bing's results. The default presumption is that any data could end up anywhere; the odd argument is that they have specially cased this and they only grab clickstream data when they have no other sources. One would expect that if the data is being collected it's being fed into the search algorithms and always being considered as a weight, not just sometimes. I hate to be a bit harsh, but that argument sounds more like spin or rationalization than a logical argument. It sounds nice but it doesn't make sense if you try to actually map it back to the real world, where somebody had to type real code to produce the effect you're hypothesizing, and they weren't writing this code with this debate in mind in advance.
- contextfree 16y agoOf course it's always being considered, but the question is how much effect it has in practice. I guess this depends on your ontological categories, but personally I distinguish between "Bing is copying (if you want to call it that, and fair enough) certain Google results" and "Bing is a copy of Google" (which I think has been alleged or insinuated, but not shown). The difference is a matter of quantity turning into quality. I think they probably have statistics that convince them of the latter, and they should show them, so we can see if people outside Google also find them convincing.
- boredguy8 16y agoI don't understand how this isn't a settled issue. If clickstream data is 1 of 1000 signals, and you create clickstream data for a specialized query that will never trigger off another signal, then your created data will be reflected. That sounds exactly like what happened. You'd have to make the argument that using this data is wrong, somehow. But to make that argument, you'd basically have to argue that users shouldn't be able to share their habits with whomever they want to. I doubt that argument can be made in a compelling way. I'm surprised Google is pushing this further, and a little disappointed.
- ddkrone 16y agoYou should be able to share your results with whoever you want but that's not what happens. You share your data with just bing and google. You don't really own any of the data you are generating so saying you should be able to share it with whoever you want is saying hop before you've jumped.
- boredguy8 16y agoWhere I click on a screen, and what I click on, is not my data?
- ddkrone 16y agoOf course not. Try inquiring google or bing about obtaining all the information they have gathered from your toolbar and see what response you get. Actually I'll tell you: "The gathered data is anonymous and we have no way of verifying who sent it to us". So how exactly do you own information that you have absolutely no access to and with no way to transfer to competing search companies?
- ugh 16y agoThat’s decidedly not what owning your clickstream data means. The argument is that users should be able to decide to give away their data, whether they can later access that data is irrelevant. Access to the collected clickstream data is simply not part of the agreement and as long as users are not coerced or tricked into agreeing that’ certainly unproblematic. (I can similarly agree to give away my photo collection without getting any right to access it ever again. That doesn’t mean that I don’t own my photo collection.)
- onecommentayear 16y agoImagine I launch a search engine with no data. Then I feed it with urls IE users click on after their google search. I will eventually end up with an exact copy of google database. So I think that this technique can be called "copying". Now if Bing uses this technique for 0.1% of their data, then it can be said that 0.1% of their data are copied from Google database.
- ddkrone 16y agoWhat do you suppose google does when they are faced with a novel query and their algorithm returns 10 equally good results as matches? At that point you might as well provide a whole bunch of users some permutation of the matches and take into account which links are clicked on the most. This is exactly what bing is doing and if you crawl the web then you will indeed end up with google's database, the only difference will be how you rank results.
- sorbus 16y agoThat's not what bing is doing in this case. What bing is doing is taking information from Google, as collected by users, and presenting that information to their own users.
- ddkrone 16y agoYa, and? They are also taking information from a whole bunch of other sites collected by their users and incorporating that information into ranking search results. I don't really see how this is inherently wrong.
- sorbus 16y agoI'm not saying that it's inherently wrong, I was attempting to explain to you how your example (Google using A/B testing to determine which results are the most relevant) had no relation to what bing is doing.
- webwright 16y ago