7 ms·
We do something similar. We precompute/aggregate exhaustively by following certain aggregation strategies. The aggregated statistics are further processed to en
by visualsearchsv 11y ago
We do something similar. We precompute/aggregate exhaustively by following certain aggregation strategies. The aggregated statistics are further processed to ensure privacy.
Differential Privacy cannot be directly applied since the underlying assumptions are too strong. An important consideration is that the error/noise added is independent of the answer. Which means that the system becomes unusable for almost all queries other than general trends.
By restricting the query structure, we no longer need large amount of noise. Privacy of hospitals and providers is also very important and cannot be encoded in the Differential Privacy framework. Again this is still a hotly debated issue. But even the most vocal supporters of differential privacy agree that it might not be directly applicable for healthcare domain.
Following are some of the paper that discuss this:
http://www.openu.ac.il/personal_sites/tamirtassa/download/conferences/anonymity_dp.pdf http://www.openu.ac.il/personal_sites/tamirtassa/download/co...
http://www.jetlaw.org/wp-content/uploads/2014/06/Bambauer_Final.pdf http://www.jetlaw.org/wp-content/uploads/2014/06/Bambauer_Fi...
- yummyfajitas 11y agoSo I'm not really an expert on privacy (my main interest in differential privacy is avoiding overfitting), but isn't differential privacy by definition necessary on an individual level? I.e., if you don't have differential privacy, then by definition there is a de-anonymizing query and you can get PII out. Privacy of hospitals/providers is a separate issue, and yeah, it's pretty clear that differential privacy doesn't work for them. Thanks for the links, I'll check them out. Edit: after reading your second article, it's deeply misleading. They assume that to compute a mean, one must compute 2 queries - sum(x) and len(x), each of which must be differentially private. But that's totally wrong! You can in fact run the query mean(x) + noise, and this last query itself can be differentially private. The article also notes that queries on small data sets require more noise to be differentially private, which is totally true, and obvious. This also, however, ignores the fact that most statistical inference drawn from such queries will be nonsense even without differential privacy. See, e.g., this article for an example of why: https://www.chrisstucchio.com/blog/2015/ab_testing_segments_and_goals.html https://www.chrisstucchio.com/blog/2015/ab_testing_segments_... This is a very bad critique of differential privacy.
- visualsearchsv 11y agoThe second article is undoubtedly flawed in several aspects, that's why I put clifton et. al. first, which I think lays out the case for studying non-formal models. Regarding """This also, however, ignores the fact that most statistical inference drawn from such queries will be nonsense even without differential privacy.""" This is not true. Just because the number returned by a count query on a very large dataset (~100 Million visits) is very small (~100) does not automatically means that the result is nonsense or can be disregarded as error. Doing that requires understanding the query and a hypothesis with good prior on expected outcome. E.g. intersection of two rare diseases. Where you would otherwise expect it to be very small, but there might be an underlying genetic reason / physiological process which might lead to higher prevalence. Or a group of hospitals using tainted batch of medicines leading to unexplained increased mortality. Consider this paper where there were only 1000 cases (only 248 strokes) per 1.6 Million patients (even larger if you consider the entire 20 Million patients present in the data). However in spite of the small number the authors showed that the increase was statistically significant by comparing with same period a year later. http://www.nejm.org/doi/full/10.1056/NEJMoa1311485 http://www.nejm.org/doi/full/10.1056/NEJMoa1311485 Again I am not denying what you wrote in the blogpost. But in medicine and the analyses for which such databases are used, the investigators have access to very good priors.
- deleted 11y ago[deleted]