5 ms·
As one of the poeple who broke the anonymity of the first Netflix Prize dataset, here's my take on why this is an issue we should worry about. First, some li
by randomwalker 17y ago
As one of the poeple who broke the anonymity of the first Netflix Prize dataset, here's my take on why this is an issue we should worry about.
First, some links:
Paul Ohm's paper on the consequences of re-identification for privacy laws and public policy: http://papers.ssrn.com/sol3/papers.cfm?abstract_id=1450006 http://papers.ssrn.com/sol3/papers.cfm?abstract_id=1450006
Ohm's article about Netflix Prize 2: http://freedom-to-tinker.com/blog/paul/netflixs-impending-still-avoidable-multi-million-dollar-privacy-blunder http://freedom-to-tinker.com/blog/paul/netflixs-impending-st...
Our FAQ about the re-identification attack on the first dataset: http://www.cs.utexas.edu/~shmat/netflix-faq.html http://www.cs.utexas.edu/~shmat/netflix-faq.html
An excerpt from our paper, showing the kind of information you can learn about a person just from movie preferences:
First, his political orientation may be revealed by his strong opinions about “Power and Terror: Noam Chomsky in Our Times” and “Fahrenheit 9/11,” and his religious views by his ratings on “Jesus of Nazareth” and “The Gospel of John.” Even though one should not make inferences solely from someone’s movie preferences, in many workplaces and social settings opinions about movies with predominantly gay themes such as “Bent” and “Queer as folk” (both present and rated in this person’s Netflix record) would be considered sensitive. In any case, it should be for the individual and not for Netflix to decide whether to reveal them publicly.
A quote from Paul Ohm on why this is problematic even if no one cares about movie privacy:
The "accretion problem" is this: once an adversary has linked two anonymized databases together, he can add the newly linked data to his collection of outside information and use it to help unlock other anonymized databases. ... Because of the accretion problem, every reidentification event, no matter how seemingly benign, brings people closer to harm. Had Narayanan and Shmatikov not been restricted by academic ethical standards (not to mention moral compunction), they might have connected people to harm themselves.
He then goes on to postulate the "database of ruin":
It is as if reidentification and the accretion problem join the data from all of the databases in the world together into one, giant, database-in-the-sky, an irresistible target for the malevolent.
At a first reading, this might all sound like science fiction, but it's a lot more plausible than most people think. On my blog http://33bits.org/ http://33bits.org/ I have several more examples of re-identification. And let's not forget that there are companies such as Acxiom and Choicepoint which already specialize in aggregating every available piece of information about people, and then selling it.
- Retric 17y agoI have legitimate access to the SSN of every person in the Army over the last several years. And plenty of other data of a similar sensitive nature linked to those SSN's. On of our more interesting problems is how to aggregate that data so a user can't extrapolate excessive information about any one person. Once you really start looking at things any information you release could be useful to an attacker. But, few people are going to take the time unless you start giving out access to a lot of detailed information.
- jonmc12 17y ago"...one, giant, database-in-the-sky, an irresistible target for the malevolent." I'm not sure, but is this such a bad thing if this giant database is open? I mean, the most incentivized and most malevolent of people can already buy my information - and they will continue to figure out how to buy more of it as it becomes available. The important thing to me personally, is that if a company has data about me, they will let me see what it is, and then allow me to correct and (in some cases) redact this information. I guess what I am saying - put that database in the sky so we can all see it. Then once the info gets tracked back to my identity, have it subject to some kind of regulation so that I can reasonably control what is available. The database should not have any trouble finding my contact information to tell me that it knows something about me.