8 ms·
An open letter to Netflix from the authors of the de-anonymization paper
- prakash 17y agofyi: randomwalker is Arvind Narayanan.
- bmickler 17y agoP.S. - BTW, when do you expect to allow your linux-using, paying customers to watch your instant streaming movies online?! (off topic rant, I know. goodbye karma!)
- nkurz 17y agoI'm saddened to see someone gloating at having helped to prevent the release of a dataset that I see as beneficial. Netflix offered an unprecedented corpus for research, and now someone is proud about helping the lawyers to lock it up. I think I just have a fundamentally different sense of privacy than the author. I think this comes out most clearly in their FAQ: Furthermore, even if the algorithm finds the "wrong" record, with high probability this record is very similar to the right record, so it still tells us a lot about the person we are looking for. So the "violation of privacy" occurs even we don't actually reveal information about the individual, even if we only provide a framework for making predictions? So if I publish a study (with backing data) that says that 38 year old males are likely to commit adultery, I've "violated the privacy" of all 38 year old males? Could someone who shares the author's worldview try to explain it? I've tried, but I just don't see it.
- randomwalker 17y ago'Gloating' is so diametrically opposed to the view expressed in the article that I have no idea how to respond. As for that part of the FAQ, it is intended as an explanation of some of the theorems proved in the paper and is a response to some of the theoretical objections we face from the data privacy community. It is not an issue that arises in practice.
- pbh 17y agoI think the parent is referring to an element of we-told-you-so in this letter. I think around the "instead, you brushed off our claims" part of the letter. It sounds a bit confrontational, though perhaps it's too late to rewrite.
- grandalf 17y agoIt sure makes the writers sound like asses.
- dasil003 17y agoThis is the part that did you in: Instead, you brushed off our claims, calling them “absolutely without merit,” among other things. Now even if you are all emotionless, 100% objective researchers with no interest other than the greater good, this very sentence will make it impossible for humans to read your letter without implying a certain level of snark. If you want to project pure motives for the letter then it would have been best to leave their reaction to your original research out of it, or, probably more appropriately, don't publish an open letter at all—contact them directly.
- madair 17y agoAh yes, Appeal to Motive, that'll do it, now I'm not going to accept anything they said because of my interpretation of their motives. http://en.wikipedia.org/wiki/Appeal_to_motive http://en.wikipedia.org/wiki/Appeal_to_motive
- dasil003 17y agoOpen letters are about politics, not logic. What an individual critical thinker such as yourself believes has no bearing.
- madair 17y agoOpen letters are about communicating a message, they are also by definition political, and it's really helpful to our social systems when they are based on logic. Analysis of motives is of course valuable. But it's not an argument against the matter at hand. It may not have been clear that I was being facetious.
- raganwald 17y ago> Could someone who shares the author's worldview try to explain it? Well, I can give you my worldview. Did you see the article "70% of HR managers turn down job candidate based on online reputation?" Here's the HN discussion: http://news.ycombinator.com/item?id=1192996 http://news.ycombinator.com/item?id=1192996 So. I rent a bunch of movies with strongly opinionated themes. I dunno, maybe every Michael Moore movie plus travel documentaries about Cuba. I also rent the entire Star Wars sextilogy and The Godfather Trilogy, which I review on IMDB.com. Some enterprising hacker ties my reviews on IMDB to my anonymized rental record. How many 47 year-old males are there in my neighborhood who watched those exact movies? The hacker publishes his "findings," of course. Now I go looking for a job and someone in HR decides that my political views are too risky, so I don't even get an interview. The author's views are that Netflix can provide the benefits you desire without compromising the privacy I desire. I think the debate should be around whether the authors are correct. If the authors are correct, then the problem isn't the authors preventing the release, the problem is Netflix failing to learn from their prior mistake. If the authors are incorrect, we should simply point that out.
- Jun8 17y agoThe important point is that just renting out the movies isn't enough to reveal your identity, you also have to review them (anonymously) on IMDB or some such site. So the solution is simple: If you rent a movie that you would object to being listed in, e.g. your public Facebook profile just don't review it. Is this so hard?
- lt 17y agoThat's incorrect. His public Star Wars reviews linked him to the Michael Moore's rentals he wished to remain private.
- Jun8 17y agoYou're right. How about no public reviews on IMDB then? Or create another different users? I rent many movies from Netflix that people would find "objectionable", this is one reason I don't create any public reviews.
- imurray 17y agoI'm saddened to see someone gloating at having helped to prevent the release of a dataset that I see as beneficial. I'd encourage people that usually skim just the comments to read the post, which was not gloating at locking up the dataset. The thrust of the post is about what a shame it is that the contest was canceled and how to make sure future contests can work. [Note one: I have nothing to do with either side. Note two: I guess there is a gloating interpretation, with the paraphrase "you ignored us, but the FTC said we were right, nah nah nah" — but this isn't a useful or constructive way to continue the conversation.]
- lvvlvv 17y ago"you ignored us, but the FTC said we were right, nah nah nah" That how I read their open letter. And my interpretation of "you should have worked with us" is as "you should have hired us as consultants"
- gwern 17y ago> So if I publish a study (with backing data) that says that 38 year old males are likely to commit adultery, I've "violated the privacy" of all 38 year old males? Yes. The difference between 'John Elks of 7 Arborview is a rapist' and 'there is a 0.5 correlation between living in Pleasantville and being a rapist' or '38 year old males commit 10% more rapes than average' etc. is solely one of degree and not kind. Suppose I have a set of datapoints like that. And let's say each datapoint applies to only half the population. How many datapoints before I have broken your privacy and linked you to the furry porn you like to rent? Well, I'm guessing you're an American male. The US population is ~300 million, and roughly half of that is male, so 150 million. The first datapoint pins you down to within 75 million. The second, down to 33 million. The third, down to 16 million, the fourth, 8 million, the fifth 4 million, the sixth 2 million, the seventh 1 million, the eighth 500,000 (starting to feel nervous yet?), the ninth 250,000, the tenth 125,000, the eleventh 75,000, the twelfth 30k, etc. until the 25th or 26th specifies just 1 - you. Now, tell me: Where in this slippery slope did it suddenly flip from not being a privacy violation of some degree, to being a privacy violation? Was it at the 5th bit of information? Are you damaged at the 12th bit of information? Or did it take until the 24th or 25th bit of information before it magically flips from being good science to bad privacy violation? Is it fine just so long as it might also be your neighbor down the street, even though most people would shun you based on far less than a 50-50 chance of things like being a child rapist? (An employer on the bench might regard a 10% chance of you being objectionable as being too much; that only requires, what, 18 bits of information?) Predictions embody a great deal of information. That's how Bayesian statistics and statistics in general work, after all.
- tedunangst 17y agoI'm very curious about what bits of information you think exist that so precisely bisect the population.
- conover 17y agoThe Electronic Frontier Foundation does it to browsers pretty easily. http://panopticlick.eff.org/ http://panopticlick.eff.org/ And according to another article by them, all you need is zip code, gender and birthday to identify someone with a high degree of certainty. http://www.eff.org/deeplinks/2009/09/what-information-personally-identifiable http://www.eff.org/deeplinks/2009/09/what-information-person...
- madair 17y agoNetflix has a lot of subscribers, but it doesn't have that many, and in particular, in your neighborhood, does it have that many who are 38 year old males who like action movies and Michael Moore? It's a question partly of scale. It's big, but not that big that we can't get dangerous results with demographics & statistical analysis. It comes back to the contemporary problem of statistics: We, programmers included as evidenced by many discussions of the sort, have a hard time getting it. We are only at the cusp of understanding what statistics can do with large data sets. Now, we can either work with researchers and organizations to deal with the very real concerns, or we can simply refuse to believe that giraffes don't exist because we never saw an animal with such a long neck before. It's the basic problem of progress.
- deleted 17y ago[deleted]
- pbh 17y agoIt seems to me that any data at all will necessarily reduce the entropy of the probability distribution of members' preferences, likes, dislikes, habits, and so on (i.e., their privacy). The authors seem to brush off the "greater good" argument here, but I don't understand how any large scale data release can happen without at least some reference to such an argument given that context. Given that, the authors seem to be making a fairly strong claim here: that no large scale "anonymised" data release should ever happen. Is that helping anyone in the context of movie viewing? And is it hurting anyone other than academic researchers, given that companies share more sensitive data anyway?
- raganwald 17y agoThe authors mentioned two alternatives to the current form of large-scale release. First, opt-in. Second, contestants submit programs that run on the anonymized data but the contestants do not have access to the data itself. Could either of these approaches contribute to the greater good without compromising privacy?
- pbh 17y agoI do not have any data regarding opt-in, but my impression was that as a rule, no one opts-in, and no one opts-out of basically anything (excepting the notorious cases, e.g., Real Player). If it worked, and somehow gave an at least somewhat unbiased sampling of the data, opt-in would obviously be best. I am completely unsatisfied with the submit-and-run model. Feature engineering seems to rely on knowing your data really intimately, and that does not seem possible in a submit-and-run model.
- smokey_the_bear 17y agoWhat happened with Real Player?
- pbh 17y agoThe Real Player installer, over the course of maybe five to ten years, was so pushy about installing extra, unwanted software and sending private data that it garnered a reputation that caused people to be really careful when installing it, if they installed it at all. It is not really a perfect example for people actually opting-out, because I think one of the many criticisms was that it was often either not possible or extremely difficult to determine how to opt-out of its features (sending titles of files being played, annoying message center and ad popups, packaged additional software). That said, check out the Wikipedia page for further details.
- stevoski 17y agoIIRC, in some countries census data is sometimes released with small, intentional errors to prevent the ability to locate specific individuals. Make a 36 year old sometimes a 37 year old or a 35 year old. Make a 180 cm person sometimes 182 cm or 178 cm. Small enough errors not to make the aggregate data invalid, but enough to make it hard to identify individuals from the data. Perhaps this is a partial solution for the Netflix dataset.
- nkurz 17y agoThis is the approach that Netflix took with the initial data. The paper referred to shows that this is insufficient, and does little to ease privacy concerns. The general problem is that if you 'fuzz' up the data enough to make identification impossible, it's no longer useful as a dataset.
- wooster 17y agoThe Census obfuscations have apparently screwed up a variety of research findings: http://freakonomics.blogs.nytimes.com/2010/02/02/can-you-trust-census-data/ http://freakonomics.blogs.nytimes.com/2010/02/02/can-you-tru...
- randomwalker 17y agoI'm a little taken aback by the tone of some of the comments here, so I thought I'd offer a few points of clarification. * For the longest time we mostly stuck to doing the math. We certainly didn't call for the contest to be cancelled, and we had nothing to do with the lawsuit. But when some people implied that we were responsible for the mess that ensued, we were kind of pulled into it. We posted this as a way of explaining our point of view and reaching out to see if there's a possibility of collaboration. * The sadness we expressed is genuine. The reason we brought up Netflix's response to our paper wasn't "snark" or "gloating." Rather, we were pointing out that the cancellation of this contest was rather needless, because if they had acknowledged the privacy risks back when we published the paper, they would have had more than enough time to deploy an opt-in system for this contest. I think it is really unfortunate that that didn't happen. * Someone wanted to know exactly what I thought of the "greater good" argument. Well, I'll tell ya. I'm vehemently opposed to it and I think it's a dangerous slippery slope. I think this point of view is enshrined in the ethos of this country -- "better let ten guilty men walk free than to convict one innocent man," etc. I don't think anyone has the moral authority to decide that the privacy concerns of a few can be sacrificed. * There is a specific reason we chose the open letter format rather than communicating with Netflix directly. Acutally, two reasons. First, there are many data privacy researchers who are at least as qualified as we are for this role. We wanted to make sure the community had the opportunity to participate in whatever ensures, rather than just us. Second, I'm sure there are many companies other than Netflix who have a similar need for privacy preserving data mining. If Netflix doesn't take us up, perhaps one of the others will. Bottom line, since there are multiple parties on both sides, and we don't really know who they are, we felt it is better to have this dialog in public. * Finally, it is understandable when something like this happens to want to find someone to blame. But think twice before shooting the messenger.
- Aron 17y agoI am gonna make a note on only a small part of this, which is that 10 guilty men is different than 'an uncountable number of guilty men'. In other words, all principals have soft margins in practice.
- earl 17y agoPlease. You directly enabled the lawsuit. At least have the integrity to acknowledge your part -- your pretense "Oh wow, a lawsuit just happened! But I had nothing to do with it!" -- is pretty stupid. I'll bet money the most likely result of this lawsuits -- and your actions are a big piece of it -- is Netflix, et al, will release no more datasets to the public. Instead, only researchers under NDA will be allowed to work with the data, as was basically the tradition before this. Good job.
- hooande 17y agoI looked at the methods used in the paper and it's clear that my definition of "privacy" varies greatly from the author's definition. Essentially they are saying that if you know what rating someone gave 8 movies and the date that they gave those ratings, you can find a sample of their rating list (or something very similar to it) with 99% accuracy. So freaking what? "Evidence" that flimsy wouldn't stand up in a local bar argument, much less a court of law. The records are still completely and totally anonymous. No names, no addresses, no way to identify anyone...nothing but a strong statistical correlation to a set of ratings in a database. It sounds like they have a problem with the power of predictive modeling and not with the handling of anonymous data. Esstentially what they are saying is "if we know a little about you, we find can out things that we didn't know with a very high degree of accuracy, but no certainty.". Duh. That's what the whole netflix prize was about...using known data to make strong predictions about unknowns. They had some interesting methods (especially in their similarity calculations) but this has nothing to do with privacy.