7 ms·
I'm the first author of this piece, happy to answer any questions. Much of my previous research on re-identification (http://33bits.org/about/ http://33bits.org
by randomwalker 12y ago
I'm the first author of this piece, happy to answer any questions. Much of my previous research on re-identification (http://33bits.org/about/ http://33bits.org/about/) has been discussed on Hacker News.
- ihaveaq999 12y agoHow does this interact with the field of differential privacy?
- randomwalker 12y agoDifferential privacy is a different way of releasing data that avoids the problems of re-identification. The theory is well developed, the tools are starting to get there, but the hardest part is that it requires a behavior change from data analysts -- you need to formulate the desired computation algorithmically instead of just poking around the data. This has proved to be a formidable barrier.
- mjw 12y ago> you need to formulate the desired computation algorithmically instead of just poking around the data. This has proved to be a formidable barrier. Very true. Furthermore, if you start to think seriously about making differential privacy guarantees across a whole company, it gets really hard and impractical. It's not enough, as you might naively hope, to certify individual data analysis jobs separately as being "differentially private". Taken cumulatively they can amount to something which isn't [1] It's not even sufficient to certify whole planned and delimited programmes of use for whole datasets as differentially private, if the company also holds (or wants to hold in future) other data on some of the same users and there's some chance of new actions being taken which directly or indirectly depend on both sources of data. For the theory to truly apply, one needs to track (in a particular formal sense) how much "privacy budget" is used up by pretty much every user-data-dependent action taken across a company and across all datasets you hold pertaining to any overlapping set of users. Once your preallocated budget is used up it's pretty much game over in terms of taking any further actions based on any data (acquired now or in future) on any of those users, unless these actions can be based on inferences which were already obtained within the original budget. These aren't necessarily insurmountable problems, and I'm sure there are new tools being developed to help which I'm not up to date with. (One which I am aware of is Microsoft's PINQ project [2]). Still, it seems the organisations with the biggest chance of actually making this work, are those whose relationships with sets of users are inherently transitory and/or firewalled off from eachother. For example a B2B company who operate one-off surveys for clients, throwing away the raw data afterwards. The fundamental problem is that differential privacy guarantees a degree of resistance to an incredibly strong adversary -- one with unlimited prior knowledge about your users, who is able to bring unlimited intelligence and computational resources to bear on drawing inferences about them based on every action you ever take conditional on user data. It's impressive that one can obtain any guarantees at all under this model, but perhaps not surprising that it can be hard to scale up and compose the guarantees without things blowing up. That's not to say that there's isn't value in trying to obtain differential privacy guarantees for smaller scale pieces of work. It's still considerably better than other more naive pseudo-anonymisation methods. [1] at least not to the same strength -- see http://en.wikipedia.org/wiki/Differential_privacy#Composability http://en.wikipedia.org/wiki/Differential_privacy#Composabil... for the less handwavey version [2] http://research.microsoft.com/en-us/projects/PINQ/ http://research.microsoft.com/en-us/projects/PINQ/
- siculars 12y agoWhat are your recommendations for anonymizing PHI data, for example? Saying there is no way to anonymize data isn't really a solution and beyond that, I don't think it's true. What practitioners need is a canonical reference or toolset that you could feed a csv file to, tell it which columns should be scrambled and out comes an anonymized data set. Yes, I realize the "tell it" part is of concern, well that could be mitigated by smarter tools and more knowledgable practitioners - knowledge gained from a canonical reference.
- randomwalker 12y agoRegardless of whether PHI can be anonymized in a fool-proof way, we can agree that careful anonymization is better than a superficial one, and so your question is valid and important. We can't automate the process (in part because the transformations necessary are much more complex than "scrambling"), but knowledgeable practitioners can go a long way. I'm knowledgeable but not a practitioner, so I'm not the best source. In #8 of our report (on the Heritage Health data), you'll notice that while I took Khaled El Emam to task for claims about quantifying risk, I do acknowledge that he did a very good job of de-identification. I don't think there's exactly a "canonical reference" (except HIPAA's superficial list of 18 identifiers), but reports written by practitioners like El Emam are probably useful documents.
- jandrewrogers 12y agoIn order to have minimally robust anonymization you must strip out all implicit or explicit references to locations and times. Unfortunately, space and time values are material to the analysis of most non-trivial data models (it certainly is for health data) so stripping that out is not really an option. Few people appreciate the robustness and generalizability of spatiotemporal coincidence analysis for reconstructing relationships in anonymized data both within and across many unrelated sources of anonymized data, even sources that are identifying entities that are not people (like anonymous vehicle tracking). There are enough anonymous entity tracking data sources available to algorithmically reconstruct relationships to most other "anonymous" data sources. I've seen it done many times and the capabilities are jaw-dropping in part because it violates human intuition as to what it is possible with such data sets.