3 ms·
One troubling re-identification attack for medical data is the trail re-identification method [0]. A lot of privacy analysis will consider the data to be in the
by aabaker99 5y ago
One troubling re-identification attack for medical data is the trail re-identification method [0]. A lot of privacy analysis will consider the data to be in the shape of a table T with some columns A,B,C and use the notation T<A,B,C> to describe a de-identified dataset. The trail method will take multiple de-identified datasets, each from a different hospital, T_1<A,B,C> T_2<B,C,D>, T_3<C,D,E> and use their shared columns to narrow down on a set of individuals.
So, even though each hospital may have a legitimately de-identified dataset in isolation, it is not de-identified when combined with the (also de-identified) data from another hospital. The risk of this attack increases as patients visit more hospitals. We humans are fairly long-lived and tend to move around so it may be substantial. (That being said some hospital systems are quite large like Kaiser Permanente and serve huge areas so visiting multiple hospitals doesn't necessarily create multiple tables.)
[0] https://dataprivacylab.org/dataprivacy/projects/trails/trails2.html https://dataprivacylab.org/dataprivacy/projects/trails/trail...
- alistairSH 5y agoFor this attack to work, wouldn't one of the tables need to contain PII of some sort? If A,B,C,D,E are all de-identified, the aggregate is still de-identified? But, if E is SSN (or some other PII data), then the entire set can be re-identified?
- taejo 5y agoThat's one option: you combine protected, de-identified information with unprotected (e.g. non-health) information to re-identify the protected information. But also, something like Facebook allows you to target a person who lives in $TOWN, works at $COMPANY, born in $YEAR, even if you don't know that person's name or SSN.