2 ms·
Most of us recognize the value of having such detailed data, and also the consequences of it (like in this article). I would prefer not to have to "trust" the
by aneesh 18y ago
Most of us recognize the value of having such detailed data, and also the consequences of it (like in this article). I would prefer not to have to "trust" the company storing the data, and I think leaks like this make it clear this can't be the approach in the future. So how can this problem be solved from a technology perspective?
I'm aware of k-anonymization, where you change the data to make it mathematically impossible to identify an individual data point more specifically than among k entries. So for example
Age Weight Disease
18 150 Cancer
45 203 Diabetes
37 197 Heart Disease
becomes
Age Weight Disease
* * Cancer
[36-45] [195-210] Diabetes
[36-45] [195-210] Heart Disease
for k=2, and it's now impossible to know for sure the disease of the 45-year-old, whereas you could deduce that information from the earlier records.
Another approach, used by some major search engines for ad-targeting, is to insert random noise into the data.
What other solutions are there? If you build software that captures sensitive data, how do you deal with it?
- pierrefar 18y agoOne set of data is not the issue. It's when you start compiling knowledge (I almost want to say "evidence") from multiple data sets that privacy is really invaded. Couple your anonymous dataset with the search histories of the people involved, and you'll likely get your answer. Or couple it to their travel info to see which hospital are they likely to have visited so regularly - does the hospital have a specialist diabetes or heart disease unit?