5 ms·
Let’s say you wanted to count how many of your online friends were dogs, while respecting the maxim that, on the Internet, nobody should know you’re a dog. To d
by quantisan 9y ago
Let’s say you wanted to count how many of your online friends were dogs, while respecting the maxim that, on the Internet, nobody should know you’re a dog. To do this, you could ask each friend to answer the question “Are you a dog?” in the following way. Each friend should flip a coin in secret, and answer the question truthfully if the coin came up heads; but, if the coin came up tails, that friend should always say “Yes” regardless. Then you could get a good estimate of the true count from the greater-than-half fraction of your friends that answered “Yes”. However, you still wouldn’t know which of your friends was a dog: each answer “Yes” would most likely be due to that friend’s coin flip coming up tails.
Source: Google's RAPPOR project
I pointed to some open source repos on my blog post from 2015 https://www.quantisan.com/a-magical-promise-of-releasing-your-data-and-keeping-everyones-privacy https://www.quantisan.com/a-magical-promise-of-releasing-you...
- cm2187 9y agoIt doesn't really protect privacy, unless your first coin is highly biased toward the random answer. You can still infer that there is a higher likelihood that this individual is a dog. A few 60-70% reliability inferences on various dog related characteristics and you can identify a dog with 95% chance.
- tempay 9y agoThis would be a flawed implementation. The method still works, you just have to be extremely careful to avoid exposing users through correlations when collecting anything more than a single value just once. The aforementioned RAPPOR paper[1] covers this in Section 6. [1] https://arxiv.org/abs/1407.6981 https://arxiv.org/abs/1407.6981
- cm2187 9y agoIf I understand correctly, they just call it a limitation of the approach and invite to sample the user base rather than collect data systematically (plausible deniability is a bit of a moot point, once you have been exposed as a dog with 95% confidence, pleading that there is still a small chance you might not be one doesn't really help). I am not sure how any of that helps in a mass collection of data like OS telemetry.
- frankmcsherry 9y agoI suspect you did not understand correctly? I just read their section 6, and they neither "just call it a limitation" (the section is three pages long, with several recommendations) nor invite anyone to sample the user base (this generally just focuses the privacy loss on the sampled people). Which text were you reading that lead you to this conclusion?
- cm2187 9y agoOn sampling: > It is likely that some attackers will aim to target specific users by isolating and analyzing reports from that user, or a small group of users that includes them. Even so, some randomly-chosen users need not fear such attacks at all... For the limitation, the whole section 6.1 explains that this only protects a single question. If you collect more than single question, you must rely on other techniques to protect the privacy.
- frankmcsherry 9y agoYes, I think you've misunderstood. The text you've quoted is about how a random subset of the population is already immune to the issue of repeated queries, not that subsampling the population helps in any way. If you don't interrupt the quotation mid-sentence, it reads: > Even so, some randomly-chosen users need not fear such attacks at all: with probability (1/2 f)^h, clients will generate a Permanent randomized response B with all 0s at the positions of set Bloom filter bits. Since these clients are not contributing any useful information to the collection process, targeting them individually by an attacker is counter-productive. The whole of section 6.1 is not about how it only protects a single question, it is about how one ensures that the single-question protections generalize to larger surveys, concluding that > This issue, however, can be mostly handled with careful collection design.
- cm2187 9y agoSampling and taking a random subset of the population are synonymous. But my point is precisely that this technique helps with a single question. As soon as you are doing continuous mass collection you don't really get any privacy protection from this technique, and you have to rely on other techniques (encryption, etc).
- mirimir 9y agoOK, thanks :) But ... [please see reply to omarforgotpwd].
- tkuichooseyou 9y agoIn that case though, you would know that the "No" friends are definitely not dogs, and the "Yes" friends are possibly a dog, so it seems like the dogs would still not be completely anonymous. Wouldn't the dogs be better off not partaking in the survey and being narrowed down into a group of possible dogs?
- tempay 9y agoThe implementation of this should give a random answer when not being truthful. If the coin comes you heads you answer truthfully. If it comes up tails, you flip the coin again and answer if the yes if the coin is heads and no if the coin is tails. You can then no longer know if anybody is (or is not) a dog. The probabilities can be adjusted to provide more or less privacy (while making the data less or more useful). For example, if you only answer truthfully 0.1% of the time it would be hard to know anything about anyone, at the cost of knowing the total number of dogs less precisely.
- tajen 9y agoThis often helps tracking people's opinions indeed. "Here's a neat trick: If you want to work out whether your favorite celebrity is a republican, just Google their name and see if they talk about politics. If the answer is no, then they're a Republican."
- 1_player 9y agoIs there a name for this algorithm?
- LaGrange 9y agoThe entire approach is called differential security, as in the headline of the thing we're commenting upon ;-) But, as someone mentioned, if the coin comes tails you should answer with another coin flip, not "yes".
- frankmcsherry 9y agoThe coin flipping approach is called "Randomized Response" and dates back to the 60s. https://en.wikipedia.org/wiki/Randomized_response https://en.wikipedia.org/wiki/Randomized_response
- pfortuny 9y agoThe proper way is: you flip a coin, if it comes up heads, you say the truth, otherwise you say whatever you want. Then you infer an estimate using Bayes' theorem. Otherwise it is not private, as a reply has pointed out.
- dpriv123 9y ago" The proper way is: you flip a coin, if it comes up heads, you say the truth, otherwise you say whatever you want." Saying "whatever you want" will incur a very large sampling error, especially as the population of those saying whatever increases. What is needed is a notion of scalable privacy, where as the population of those saying "whatever" increases the privacy strength also increases yet the absolute error remains at worst constant. https://arxiv.org/abs/1708.01884 https://arxiv.org/abs/1708.01884
- pfortuny 9y agoMmmhhh... I was trying to ELI5. I understand the sampling error may be large but I cannot see the inherent problem. Could you explain please? (I mean, what do I have to do if I get tails? or do we change the coin). Honest question, just too lazy to read the manuscript you link...
- dpriv123 9y agocopying verbatim relevant sections " the estimation error quickly increases with the population size due to the underlying truthful distribution distortion. For example, say we are interested in how many vehicles are at a popular stretch of the highway. Say we configure flip1 = 0.85 and flip2 = 0.3. We query 10,000 vehicles asking for their current location and only 100 vehicles are at the particular area we are interested in (i.e., 1% of the population truthfully responds “Yes"). The standard deviation due to the privacy noise will be 21 which is slightly tolerable. However, a query over one million vehicles (now only 0.01% of the population truthfully responds “Yes") will incur a standard deviation of 212. The estimate of the ground truth (100) will incur a large absolute error when the aggregated privatized responses are two or even three standard deviations (i.e., 95% or 99% of the time) away from the expected value, as the mechanism subtracts only the expected value of the noise." "In this paper, our goal is to achieve the notion of scalable privacy. That is, as the population increases the privacy should strengthen. Additionally, the absolute error should remain at worst constant. For example, suppose we are interested in understanding a link between eating red meat and heart disease. We start by querying a small population of say 100 and ask “Do you eat red meat and have heart disease?". Suppose 85 truthfully respond “Yes". If we know that someone participated in this particular study, we can reasonably infer they eat red meat and have heart disease regardless the answer. Thus, it is difficult to maintain privacy when the majority of the population truthfully responds “Yes". Querying a larger and diverse population would protect the data owners that eat red meat and have heart disease. Let’s say we query a population of 100,000 and it turns out that 99.9% of the population is vegetarian. In this case, the vegetarians blend with and provide privacy protection of the red meat eaters. However, we must be careful when performing estimation of a minority population to ensure the sampling error does not destroy the underlying estimate."