3 ms·
Intuitively, it would make sense to have a 1/n fraction: you must be one out of n people in the database, right? However, I can think of two counter arguments (
by itcrowd 7y ago
Intuitively, it would make sense to have a 1/n fraction: you must be one out of n people in the database, right? However, I can think of two counter arguments (correct me if I am wrong).
1) Consider that your date of birth + postcode appears twice in a dataset. Naively, the operator can think to correctly identify you 50% of the time (you are either entry #1 or entry #2). But, it is also possible that you are not actually in the list. Maybe there are two other people in the area with your DoB which are not in the database, thereby reducing the odds of guessing correctly. You have to correct for this sampling bias. Also, the result is now not necessarily in the form of 1/n anymore.
2) Consider that you now have two datasets that together contain the whole population (but there is overlap between the datasets that is difficult / impossible to remove). Your DoB + postcode appears 3 times in the first set and 5 times in the second dataset. Then ask again: what are the odds of identifying you from the 8 candidates? Again, this does not give a solution in the form of 1/n.