4 ms·
I think the AOL search leak reference is this: https://en.m.wikipedia.org/wiki/AOL_search_data_leak https://en.m.wikipedia.org/wiki/AOL_search_data_leak Even i
by ceras 5y ago
I think the AOL search leak reference is this: https://en.m.wikipedia.org/wiki/AOL_search_data_leak https://en.m.wikipedia.org/wiki/AOL_search_data_leak
Even if you take more care than AOL did in anonymizing your data, the unfortunate reality is that any publication of data increases the knowledge an adversary has at identifying somebody. Anonymizing is more about reducing the chance someone is identified than guaranteeing they never will be. And high dimensional data is particularly hard to do so in a way that retains the data's usefulness.
33 bits is an old defunct blog on this topic, but it has some interesting posts and academic papers if you want to go down the rabbit hole: https://33bits.wordpress.com/about/ https://33bits.wordpress.com/about/
Specific paper on Netflix deanonymization: https://33bits.wordpress.com/about/netflix-paper-home-page/ https://33bits.wordpress.com/about/netflix-paper-home-page/
- jacquesm 5y agoThey key is to combine more than one anonymized dataset, this vastly increases your chances at de-anonymization. This paper is a very good starting point: https://papers.ssrn.com/sol3/papers.cfm?abstract_id=1450006 https://papers.ssrn.com/sol3/papers.cfm?abstract_id=1450006