4 ms·
> I feel that this is a net loss to society Is this just a random feeling or is there something to back this up? How are you even assessing things like the va
by NotAtWork 12y ago
> I feel that this is a net loss to society
Is this just a random feeling or is there something to back this up?
How are you even assessing things like the value of the difference in precision caused by doing something like creating (mostly) distribution equivalent (but fake) data and the cost of leaked person details to hundreds of thousands or millions of people?
- nkurz 12y agoI may be wrong, but it's more than just a random feeling. I should also point out that I'm not suggesting that all data should always be released. And also I'll reiterate that it's not just 'randomwalker' I'm referring to here. I'm disappointed that the paper he's rebutting seems to be making the silly claim that 'it's safe' instead of 'it's a worthwhile risk'. The case I'm most familiar with is the Netflix Prize. I think I can safely say (in my semi-professional opinion) that a lot of good research was published as a result of Netflix's decision to release the data that they did: http://scholar.google.com/scholar?q=netflix http://scholar.google.com/scholar?q=netflix I view these publications as a public good. It's possible the techniques described were previously known privately, but before the contest there was no description of them in the available literature. At the conclusion of the contest, Netflix briefly announced that they would have a second contest, which would involve the release of another data set. In large part as a result of the press coverage of randomwalker's work, this potential release was cancelled: https://freedom-to-tinker.com/blog/paul/netflix-cancels-netflix-prize-2/ https://freedom-to-tinker.com/blog/paul/netflix-cancels-netf... It's disputable, but I feel this additional data set would have generated a similar number of good publications, inspired new research, spread knowledge, and offered a similar benefit to the public. It can be argued that preventing further data releases has prevented further harm to the public. But I think it's important to look at the actual harm done by the release of the first dataset. While there was some degree of potential for harm, to my knowledge no individual was actually discriminated against as a result of being identified in the data set. By contrast, I think it is accepted that numerous individuals have been harmed by other online data breaches. If more attention is paid to anonymous datasets than to other matters that cause actual harm, the attention is likely misguided. I'd draw a parallel with the increase in airline security post-9/11. It's possible that confiscating oversize toiletry items and removing shoes has prevented further harm. It's indisputable (I think) that the additional hassle has had a negative cost to the public. And if more people choose to drive than to fly, the net effect is likely negative. But at least in the case of airline security, there is a clear case to be made of actual harm. Whether the current policies are a good compromise of inconvenience and safety depends on both the risk and the reward. I know less about the other cases, both for benefit and harm. Have New York taxi drivers suffered harm as a result of the poor attempts at anonymizing the data? Is there public benefit to the data that was released? Have individuals been harmed by the semi-anonymized release of health records? Was this offset by any positive effects? This is the discussion I think we should be having, rather than simply pointing at the abstract potential for harm.