6 ms·
Differential privacy is a different way of releasing data that avoids the problems of re-identification. The theory is well developed, the tools are starting to
by randomwalker 12y ago
Differential privacy is a different way of releasing data that avoids the problems of re-identification. The theory is well developed, the tools are starting to get there, but the hardest part is that it requires a behavior change from data analysts -- you need to formulate the desired computation algorithmically instead of just poking around the data. This has proved to be a formidable barrier.
- mjw 12y ago> you need to formulate the desired computation algorithmically instead of just poking around the data. This has proved to be a formidable barrier. Very true. Furthermore, if you start to think seriously about making differential privacy guarantees across a whole company, it gets really hard and impractical. It's not enough, as you might naively hope, to certify individual data analysis jobs separately as being "differentially private". Taken cumulatively they can amount to something which isn't [1] It's not even sufficient to certify whole planned and delimited programmes of use for whole datasets as differentially private, if the company also holds (or wants to hold in future) other data on some of the same users and there's some chance of new actions being taken which directly or indirectly depend on both sources of data. For the theory to truly apply, one needs to track (in a particular formal sense) how much "privacy budget" is used up by pretty much every user-data-dependent action taken across a company and across all datasets you hold pertaining to any overlapping set of users. Once your preallocated budget is used up it's pretty much game over in terms of taking any further actions based on any data (acquired now or in future) on any of those users, unless these actions can be based on inferences which were already obtained within the original budget. These aren't necessarily insurmountable problems, and I'm sure there are new tools being developed to help which I'm not up to date with. (One which I am aware of is Microsoft's PINQ project [2]). Still, it seems the organisations with the biggest chance of actually making this work, are those whose relationships with sets of users are inherently transitory and/or firewalled off from eachother. For example a B2B company who operate one-off surveys for clients, throwing away the raw data afterwards. The fundamental problem is that differential privacy guarantees a degree of resistance to an incredibly strong adversary -- one with unlimited prior knowledge about your users, who is able to bring unlimited intelligence and computational resources to bear on drawing inferences about them based on every action you ever take conditional on user data. It's impressive that one can obtain any guarantees at all under this model, but perhaps not surprising that it can be hard to scale up and compose the guarantees without things blowing up. That's not to say that there's isn't value in trying to obtain differential privacy guarantees for smaller scale pieces of work. It's still considerably better than other more naive pseudo-anonymisation methods. [1] at least not to the same strength -- see http://en.wikipedia.org/wiki/Differential_privacy#Composability http://en.wikipedia.org/wiki/Differential_privacy#Composabil... for the less handwavey version [2] http://research.microsoft.com/en-us/projects/PINQ/ http://research.microsoft.com/en-us/projects/PINQ/