6 ms·
For data to be anonymous under GDPR, it is not enough that individuals cannot be identified from the anonymized data set. If individuals can be identified when
by privacylawthrow 6y ago
For data to be anonymous under GDPR, it is not enough that individuals cannot be identified from the anonymized data set. If individuals can be identified when the anonymous data set is compared with the source data set, the anonymized data is not "anonymous".
For data to be truly anonymous under GDPR. there must be no other additional data that would allow for reidentification. If there is any other data that, when combined with the anonymous data, allows for reidentification, the data set is only pseudonymous and must be treated as personal data under GDPR.
- 3np 6y agoThat's the most concise and clear formulation of that I've seen so far. Thanks.
- smarx007 6y agohttps://en.wikipedia.org/wiki/K-anonymity#Methods_for_k-anonymization https://en.wikipedia.org/wiki/K-anonymity#Methods_for_k-anon...
- La1n 6y agok-anonymity is often only applied to "pseudoidentifiers", if you have the original dataset it'd be trivial to reverse k-anonymity applied that way. For example someone's blood pressure isn't considered an identifying variable, and would not need to be anonymised (should not too, to keep data utility high), however this would make linking against the original dataset trivial.
- smarx007 6y agoYou are right, time series data like BPM over time does not lend itself to anonymization nicely, the provider most likely will have to ask the user organizations what kind of measures (features) they need and return an average (if that's what the receiving organisation was after) that itself can be k-anonymized.
- ska 6y agoAveraged time series are very different than individual ones. This is a deep problem; it's basically unavoidable in e.g. medical research - the very factors you want to study may well be potentially identifying. The only way to address this is to balance the potential utility of the research against the potential impact of the information.
- m0nster 6y agoIn my experience, this is a question of interpretation (see e.g. Recital 26 and the question of what is "reasonably likely"). You can ask ten different experts, and you will get ten different opinions. Unfortunately, many aspects of the GDPR are interpreted very heterogeneously, both in individual countries and by different supervisory authorities within the countries themselves. For this reason, it is essential that more specific guidelines and certifications are developed for the use of different technologies, including anonymization.
- privacylawthrow 6y ago> In my experience, this is a question of interpretation (see e.g. Recital 26 and the question of what is "reasonably likely"). This is absolutely true. The hard part is that was it "reasonably likely" changes as technology changes. It's entirely possible that a data set that qualifies as anonymous today will not be anonymous in 5 years. Organizations are responsible for the data they publish. If data loses its anonymity in the future due to release of other data sets and/or improved technology, the organization releasing the data will be responsible for the release of personal data, even if it wasn't personal data at the time of release.
- m0nster 6y agoTrue. For this reason, even anonymous data can usually not be shared as open data. You have to control the environment in which the data is used to control what is "reasonably likely" (see also comment by La1n above).
- La1n 6y agoAlso this interpretation would completely block any sharing within the pharmaceutical field, where the original data is required by law to be kept for a minimum of 25 years. I personally like the definitions from UKAN, which talk about anonymous data as relating to data environments. edit: https://msrbcel.files.wordpress.com/2020/11/adf-2nd-edition-1.pdf https://msrbcel.files.wordpress.com/2020/11/adf-2nd-edition-...
- 6y ago
- Mary-Jane 6y agoCan someone explain the point of this requirement? If a malicious actor has access to the source data there's no need to compare it to anonymized data. What am I missing?
- user837382991 6y agoPrivacy and security are not the same. Security is to protect against malicious actors. Privacy is to protect data from everyone that’s not the person PII itself.
- lmkg 6y agoIf someone makes inferences on the de-identified data, or joins it against another dataset. The source dataset lets those inferences or joins be tied back to the original identifying data. The main point is that de-identified data can still be "personal" so it's regulated. If you share or make public psuedonymous data, that data is still covered by GDPR so you have to inform the individuals, have a legal basis (such as consent), let them opt out (if applicable), etc. Even if it's been pseudonymized, I would want to know if/when my data is sold to a marketing firm or whatever.
- MaxBarraclough 6y ago> The source dataset lets those inferences or joins be tied back to the original identifying data. But if the attacker lacks the source dataset, they can't do this, and if they possess the source dataset, they'd use it for their analysis rather than using the anonymised dataset.
- warkdarrior 6y agoThe point is that if the attacker can connect your user record in the source data with user # 188da24a7789d in the "anonymized" data, they can use that de-identify all information derived or built on the "anonymized" data. Oh, there is Netflix account for user # 188da24a7789d and the IRS released tax summaries for user # 188da24a7789d? That's interesting, since I know that user # 188da24a7789d is really MaxBarraclough.
- antman 6y agoAlthough what you say makes sense, which is the respective GDPR rule? I don’t recall seeing something like this.
- jacobr1 6y agoGoogle's Pair group has a great explainer here: https://pair.withgoogle.com/explorables/anonymization/ https://pair.withgoogle.com/explorables/anonymization/