3 ms·
The second article is undoubtedly flawed in several aspects, that's why I put clifton et. al. first, which I think lays out the case for studying non-formal mod
by visualsearchsv 11y ago
The second article is undoubtedly flawed in several aspects, that's why I put clifton et. al. first, which I think lays out the case for studying non-formal models.
Regarding
"""This also, however, ignores the fact that most statistical inference drawn from such queries will be nonsense even without differential privacy."""
This is not true. Just because the number returned by a count query on a very large dataset (~100 Million visits) is very small (~100) does not automatically means that the result is nonsense or can be disregarded as error. Doing that requires understanding the query and a hypothesis with good prior on expected outcome. E.g. intersection of two rare diseases. Where you would otherwise expect it to be very small, but there might be an underlying genetic reason / physiological process which might lead to higher prevalence. Or a group of hospitals using tainted batch of medicines leading to unexplained increased mortality.
Consider this paper where there were only 1000 cases (only 248 strokes) per 1.6 Million patients (even larger if you consider the entire 20 Million patients present in the data). However in spite of the small number the authors showed that the increase was statistically significant by comparing with same period a year later.
http://www.nejm.org/doi/full/10.1056/NEJMoa1311485 http://www.nejm.org/doi/full/10.1056/NEJMoa1311485
Again I am not denying what you wrote in the blogpost. But in medicine and the analyses for which such databases are used, the investigators have access to very good priors.