3 ms·
There are several metrics in addition to the k-annonimity (https://en.wikipedia.org/wiki/K-anonymity https://en.wikipedia.org/wiki/K-anonymity) that has been g
by anon1253 7y ago
There are several metrics in addition to the k-annonimity (https://en.wikipedia.org/wiki/K-anonymity https://en.wikipedia.org/wiki/K-anonymity) that has been going around the past few days. For example l-diversity (https://en.wikipedia.org/wiki/L-diversity https://en.wikipedia.org/wiki/L-diversity) and t-closeness (https://en.wikipedia.org/wiki/T-closeness https://en.wikipedia.org/wiki/T-closeness)
t-closeness, is defined as: An equivalence class is said to have t-closeness if the distance between the distribution of a sensitive attribute in this class and the distribution of the attribute in the whole table is no more than a threshold. A table is said to have t-closeness if all equivalence classes have t-closeness.
In short: the distribution of a particular sensitive value should not be further away than a distance t from the overall distribution.
Using the t-closeness metric circumvents issues associated with k-anonymity and ℓ-diversity. Briefly, k-anonymity states that a certain attribute class should be present in at least k records, which introduces ambiguity in the data set. However, if each of the k equivalence classes are the same, properties could still be resolved simply by elimination. The ℓ-diversity metric circumvents this problem by adding a further requirement: in addition to the class to being seen in k records, these records must have at least ℓ ‘well represented’ values. But if an attacker knows the real-world distribution of values, then attributes could still be disclosed with a certain probability, simply by combining different data sources