3 ms·
Consider the following scenario. A chemical X is spilled in river. A sequence of very small towns with similar populations dot the river both up-stream down-str
by throwawayjava 8y ago
Consider the following scenario. A chemical X is spilled in river. A sequence of very small towns with similar populations dot the river both up-stream down-stream. This is a dream setting for good science -- a true natural experiment with lots of genetic and environmental similarity between the groups in each small town, but radically different levels of exposure to X. Fast forward 10 years and researchers find that concentrations of X above some threshold N drastically increase the prevalence of birth defects. However, due to the size of the populations, researchers may not be able to publish their data and analysis publicly without de facto publishing almost complete de-anonymizable data.
The scenario is not so hypothetical. The circumstances and particulars change, but this situation describes a lot of medical and environmental research.
Should the EPA take the researcher's word? Maybe not. But there's a lot of workable middle ground in-between "blind trust" and "anyone with access to a library computer and a stats textbook can figure out participant 494 -- who math tells us can only possibly be Johnny D. of Smallsville, VA -- has testicular cancer".
And the situation is even more difficult than this because de-anonymization is an active area of research with some surprisingly powerful results published in recent years; i.e., it is highly non-trivial to determine if a dataset is de-anonymizable. And the answer always depends on what data about a person's identity is publicly available, which especially in the age of annual massive data leaks change from year to year. So even if researchers don't think it's possible to de-anonymize study participants using today's mathematical tools and today's public datasets, they'll often choose to keep the dataset private just in case.