4 ms·
> It's very hard to believe it is not possible to publish medical data in an anonymized fashion that does not infringe privacy rights Then you're one of today'
by throwawayjava 8y ago
> It's very hard to believe it is not possible to publish medical data in an anonymized fashion that does not infringe privacy rights
Then you're one of today's lucky 10,000: "Robust De-anonymization of Large Sparse Datasets" http://www.cs.utexas.edu/~shmat/shmat_oak08netflix.pdf http://www.cs.utexas.edu/~shmat/shmat_oak08netflix.pdf
A lot of environmental medical research results in dense and small data instead of sparse and large data, which obviously means it's even easier to de-anonymize.
And the really important take away from the above paper is that you need to know about all the other public data in the universe in order to really ensure anonymity. Even when anonymity seems like it's "obviously" not going to be a problem.
You never know what dataset will be leaked/published tomorrow correlating columns D-F in the spreadsheet you gave EPA to someone's first and last name. So privacy is like computer security: you can never be totally sure. So when the stakes are high you should be extra cautious. And the stakes when sharing medical research are very often high.
> If that's true, how people validate any medical research at all?.. I have very hard time believing that's the only way to do modern medical research known to our civilization.
This isn't exactly a new problem, so scientists have had a little while to figure out how to make peer review work.
You share the data. You share with trusted, relevant third parties who promise to preserve the anonymity of study participants. Just becuase you do not publish possibly de-anonymizable and definitely sensitive medical information on a public website does not mean you don't share the data with interested and relevant third parties.
> But you seem to have arrived to pre-determined conclusion already, and there in fact can be no middle ground that would be acceptable to you, as it seems.
On the contrary, I do literally exactly that in my very first post in this thread: There are some common-sense workarounds. For example, requiring an EPA or third-party audit of the original dataset and analysis. Or simply carving out exceptions for research whenever it can be demonstrated that an IRB (and therefore federal law) would not have allowed the research without restricting public access to data/analysis.
The situation where IRB says "no" to public disclosure should explicitly trigger exemption, with the duty falling back to the agency to prove the IRB was wrong in the first place. The fact that this exemption is not automatic is what makes this a catch-22 that pins researchers in-between their IRB and the EPA administrator in charge of evaluating their exception.
- smsm42 8y agoNothing in that article (with which I of course was familiar) says anything about "not possible to publish medical data in an anonymized fashion that does not infringe privacy rights". What it talks about is that some data sets may be improperly anonymized, and Netflix data set is one of them. It is undoubtedly true that anonymization is harder than it appears to a naive observer. It does not mean it is impossible. > important take away from the above paper is that you need to know about all the other public data in the universe in order to really ensure anonymity. I do not see how it's takeway from that paper. > You never know what dataset will be leaked/published tomorrow correlating columns D-F in the spreadsheet you gave EPA to someone's first and last name. You also never know whether somebody doesn't just break into your system, downloads all secret data and publishes them. Yes, you have no guarantee against bad actors and mistakes. But if that's the border condition then no research should be performed at all - after all, once the data exists, there's no guarantee it won't somehow get leaked out. > The situation where IRB says "no" to public disclosure should explicitly trigger exemption IRB is a part of the institution, right? So you essentially delegating the decision about the exception to (part of) the institution. What the point of having the rule then if every organization can make their own exceptions on demand? And that makes trivially easy to hide the data in sloppy research - just sprinkle some private data on it, show it to IRB, they cry out "no, this can't be released, no way!" - and voila, you are safe from review.
- throwawayjava 8y ago>..."not possible to publish medical data in an anonymized fashion that does not infringe privacy rights". 1. Sometimes it is. 2. I already gave substantive reasons why this will commonly occur in environmental medical research -- dense datasets about small populations that capture sensitive information about individuals. > I do not see how it's takeway from that paper. Because that's literally the only assumption in the adversarial model other than "access to the published dataset". And correlating with an auxiliary dataset is literally how all actual instantiations of this attack work. > just sprinkle some private data on it, show it to IRB, they cry out "no, this can't be released, no way!" - and voila, you are safe from review. 1. My proposal doesn't say "IRB happened to say you can't publicly publish this one particular protocol." My proposal says that IRB explicitly judges that no such protocol exists. 2. IRBs are not that arbitrary and are themselves audited. 3. If you're this paranoid about intent, then you can't trust the research anyways. If the research is committing intentional fraud, they could just do it at the data collection step. The way I see it, either: A. Even if we can't trust individual scientists, we can more-or-less trust a broad subset of scientists with different and well-aligned motives (as in, a whole group of: the IRB board, the PI, the people reviewing the IRB board) when they say that data can't be shared for privacy reasons; or, B. We should expect that literally dozens of scientists, all with different motives, will routinely and actively collude to commit wide-spread fraud. If (A), my approach makes sense. If we max out paranoia and go with (B), then I don't know why we should trust the data regardless of whether it's published publicly...