4 ms·
The US Census famously used differential privacy in their most recent surveys. They have a quite extensive analysis of the tradeoffs they had to make to ensure
by jackpirate 7y ago
The US Census famously used differential privacy in their most recent surveys. They have a quite extensive analysis of the tradeoffs they had to make to ensure everyone's privacy, and there is a fairly large body of academic work that analyzes (e.g.) the economic impact of this privacy preservation. The general consensus AFAIK is that they did a great job.
See https://arxiv.org/abs/1809.02201 https://arxiv.org/abs/1809.02201 for the main paper from the Census.
- shakna 7y agoI'm not sure that this can be done with certain datasets, like health data. If you had access to anonymised health data from my nation, for example, picking me personally out of the records would be extremely trivial, using just two data points: + I have an illness that only 1% of people have. + I lost my spleen whilst I was in primary school. Both of those things are actually a matter of public record, thanks to being mentioned in various local newspapers, so it's reasonable to assume someone somewhere has that data. I believe the general consensus on health data is that you have the age a person was at an incident, and the nature of an incident, you only need to have two or three incidents in your database to de-anonymise their records. The only way to combat it is to not provide the detail required for the analysis that is exactly what the above organisations wish to do. No profiles, no fine-grained demographics. And yet, people with rarer illnesses or events than mine will still stand out, so you also need to eliminate them from the dataset, even though they may well be the ones who could benefit most from this sort of widescale analysis.