3 ms·
Hi soumyadeb, It is true that there needs to be a way to store the aggregate queries "per user". In the current approach of collection this place happens to be
by kkm 7y ago
Hi soumyadeb,
It is true that there needs to be a way to store the aggregate queries "per user". In the current approach of collection this place happens to be on Server, but aggregation per user can easily be done on client-side by leveraging Browser storage.
Approaches like k-anonymity or L-diversity are good but they tackle the problem from a different perspective - making sensitive data available for querying without revealing actual content. The approach suggested in this article talks about methodology which removes the need to collect such data in the first place.
You can also check our paper: https://www.0x65.dev/pages/dissemination-cliqz.html#GreenAnalytics https://www.0x65.dev/pages/dissemination-cliqz.html#GreenAna...
We will talk in detail about this methodology- Human Web and how we use to collect data for our search, without compromising users privacy and based on client-side aggregation.
Disclaimer: I work with Cliqz.
- soumyadeb 7y agoThanks for the reference. I quickly browsed through the paper. The problem you mention seems to arise from the fact that GA is able to tie the user across different domains (about.me, depressionforum.org etc) using a shared cookie etc. Is it still an issue if 3rd party cookies are blocked and the GA is forced to set a first-party cookie. In that case, the ID GA will get for about.me would be different from depressionforum.org? Assuming the owners for these websites are different and only care about their respective stats (depressionforum doesn't care about about.me analytics), why do we need 3rd party cookies? Am I missing something here? Well, it can probably fingerprint the browser but there is no reason to store that information.
- soumyadeb 7y agoOK got it. You guys @ Cliqz want aggregate stats across websites so GA is not the right example. Depending on the aggregates you want, client-side aggregation may still be leaking privacy. You would probably need to implement differential privacy on top before you store the data.
- solso 7y agoOf course client-side aggregation can still leak privacy, it needs to guarantee that there are no explicit or implicit elements that would allow record-linkage on the server-side. On the diff. privacy, note that you mention before you store the data, the problem is not only there, we want to prevent the data to be send at all. Diff privacy on the client-side is tricky if distributions are unknown. That's why we go for a "simpler" approach, all records send by any user should be unlikable, always. If aggregation is needed to satisfy the use-case, it can only be done on the client itself. Re-identifiability only becomes possible if a mistake is done, no a priori distributions are needed. IMHO this approach is easier than diff. privacy at the cost of being less expressive. Diff. privacy data allow for multiple use-cases where we do not (as all records) have to be independent from one another. I'm sure it's obvious that I work at Cliqz and on this very topic :-) Tomorrow there is a more technical article about our data collection, hope you like it. Also, I would like to add that any methodology applied to protect the data of the user is welcome, does not have to be ours at all. There is one caveat though, the privacy protection has to be on origin, a solution that send data that then has to be anonymized is no good in our book; because there is no guarantee that the raw data is removed. For our use-case we believe it was the easier way to get the data we needed while respecting the privacy of the user.