4 ms·
Differential privacy appears regularly on Hacker News, with either theoretical articles or projects that aim to implement it. Yet often there is a huge gap betw
by Cynddl 9y ago
Differential privacy appears regularly on Hacker News, with either theoretical articles or projects that aim to implement it. Yet often there is a huge gap between both.
For example, Apple has touted about its use of differential privacy, but researchers [1] have shown that the privacy budget is reset every day and the parameters, buried inside the code, lack a proper derivation. Similarly, Uber seems to use DP for internal analytics. However, the proposed model does not seem really robust and does not provide accurate results at all [2]. One should always carefully review claims associated with implementations of differential privacy.
[1] https://arxiv.org/abs/1709.02753 https://arxiv.org/abs/1709.02753
[2] https://github.com/frankmcsherry/blog/blob/master/posts/2018-02-25.md https://github.com/frankmcsherry/blog/blob/master/posts/2018...
- joker3 9y agoIt's like the early days of cryptography. Everybody was rolling their own algorithms because no one realized how hard it is to do that properly. Eventually we all wised up. I'm hopeful that DP will follow a similar path.
- singhrac 9y agoOne problem I've found with differential privacy is that no one talks about how to set \epsilon. I've read this book, and it's quite well written and complete, but as the title says it focuses on the algorithmic foundations. This paper [1] is much better for practitioners, and actually gives very reasonable values for the privacy guarantee (e.g., (1.2, 1e-9)), and builds on this great paper: [2]. Worth a read if you train neural networks. [1]: https://arxiv.org/pdf/1710.06963.pdf https://arxiv.org/pdf/1710.06963.pdf [2]: https://arxiv.org/pdf/1607.00133.pdf https://arxiv.org/pdf/1607.00133.pdf
- wackspurt 9y agoBased on my limited understanding* of differential privacy, it falls short on exactness (of aggregate values) and robustness (against malicious clients). I've lately been studying the literature on function secret sharing and I think it is a better alternative to DP. Take this paper: https://www.henrycg.com/files/academic/pres/nsdi17prio-slides.pdf https://www.henrycg.com/files/academic/pres/nsdi17prio-slide... Prio: Private, Robust and Scalable Computation of Aggregate Statistics Data collection and aggregation is performed by multiple servers. Every user splits up her response into multiple shares and sends one share to each server. I've understood how private sums can be computed. Let me explain it with a straw-man scheme. Example (slide 26): x_1 (user 1 is on Bay Bridge):- true == 1 == 15 + (-12) + (-2) x_2 (user 2 is on Bay Bridge):- false == 1 == (-10) + 7 + 3 ... If all users send shares of their data to the servers in this manner AND as long as at least one server doesn't reveal the identities of the people who sent it responses, the servers can exchange the sum of the shares they've received. Adding the three responses will allow the servers to infer that there are _ number of users on Bay Bridge without revealing their identities. This system can be made robust by using Secret-shared non-interactive proofs (SNIPs). This allows servers to test if Valid(X) holds without leaking X. The authors also bring up the literature on computing interesting aggregates using private sums: average, variance, most popular (approx.), min and max (approx.), quality of regression model R^2, least-squares regression, stochastic gradient descent. Bottom line: I found the discussion on deployment scenarios very interesting. Data servers with jurisdictional/geographical diversity, app store-app developer collaborations for eliminating risk in telemetry data analysis, enterprises contracting with external auditors for analyzing customer data, etc. * - I understand the randomized response and, to some extent, the RAPPOR technique (used for collecting Chrome telemetry data) but the other literature in that community goes over my head. * * - This technique is a black box to me at the moment.
- majos 9y agoApple's epsilon reset problem is real, but it's worth pointing out that they use additional heuristics based on hashing that plausibly add another layer of privacy [1]. Plausibly, not provably, but it's a bit more than just resetting epsilon. I believe Google and Microsoft use similar tweaked forms of differential privacy. In particular, note that all of these companies -- again, going off public papers -- use the "local" variant of differential privacy, which requires less trust on the user's part. The question of "lifetime" differential privacy, for a single user across different computations and datasets, is still fairly open as far as I know. [1] https://machinelearning.apple.com/docs/learning-with-privacy-at-scale/appledifferentialprivacysystem.pdf https://machinelearning.apple.com/docs/learning-with-privacy...