3 ms·
This article fails to take into account that time and time again we've seen that 'anonyimized aggregate data' is never truly anonymous. In the AOL anonimized d
by AdamSC1 10y ago
This article fails to take into account that time and time again we've seen that 'anonyimized aggregate data' is never truly anonymous.
In the AOL anonimized data leak there were plenty of individuals identified:
http://www.nytimes.com/2006/08/09/technology/09aol.html http://www.nytimes.com/2006/08/09/technology/09aol.html
MIT researchers also showed that four anonymous purchases are enough metadata to identify 90% of individuals:
https://www.technologyreview.com/s/536501/data-sets-not-so-anonymous/ https://www.technologyreview.com/s/536501/data-sets-not-so-a...
And, a personal favorite of mine where researchers from Standford and Princeton are reporting at the World Wide Web Conference this April: "Researchers found that they could identify the person behind an 'anonymized' data set 70% of the time just by comparing their browsing data to [often public] social media activities"
https://www.techdirt.com/articles/20170123/08125136548/one-more-time-with-feeling-anonymized-user-data-not-really-anonymous.shtml https://www.techdirt.com/articles/20170123/08125136548/one-m...
It would not be hard to buy a zipcode worth of data and compare it to known facts about a person until you de-anonymized it.
- jt2190 10y agoWhile I agree with the sentiment that companies have a poor track record of protecting personal data, I'd like to clarify that a "zip code of data" is not anonymized intentionally, which is what makes picking out individuals so easy. There is a technique of intentionally anonymizing data [1] that I learned about because that Apple was talking it up in relation to storing health information. I'm only a layman, but my understanding is that it makes it much, much harder to do an analysis like you describe. [1] The term is Differential Privacy. Here's a tutorial video (1 hour, 34 mins): https://youtu.be/ekIL65D0R3o https://youtu.be/ekIL65D0R3o
- A_Crazy_Idea 10y agoSo you learned about an technique. That's nice. I'm flying a kite with two strings. Certainly big data firms use 3.
- AdamSC1 10y agoThe zip code was intentionally not anonymized because it was seen as acceptable to release, and that is part of the issue of anonymized data. US Zip codes serve around 7,500 per code on average (US Census 2010) which is different than common wealth postal codes like in Canada where they serve an average of 19 households (25 - 75 people). But, zip data can be interchanged with plenty of other unique identifiers on the web. Maybe it is browser language setting, or version of java etc. Think of it like an Excel spread sheet, if in column A you can have options "1" or "2" then in a list of 100 people there will be at least 50 who share the same data footprint. If you keep adding columns from B onward with the same logic eventually you'll have pretty unique strings. Things like searching history or web history are even worse. Ever done a search for a pizza place near your address, or Google map directions from your home to another location? That identifies you pretty easily. So does connecting to your works website, and the school your kid attends. Web browsing data is nearly impossible to anonymize by its nature unless it was compiled to something like "XX% of users in Zip XXXXXX visited website.com" As for differential privacy, it is a nice emerging theory, but there are challenges with it as there is a significant trade off right now in terms of data accuracy when applying differential privacy. It is primarily effective at casting doubt on if variable "A" about user "B" in a data set is true or not, but if you don't have a specific target or specific metric then enough of the data is true that it could still in theory be deanonymized, and since most of the anonymity is based on incorrect variables in a data-set, all it would take to reverse engineer it is a large enough data-set and a few known variables. I hope people like Apple continue to champion the advancement of differential privacy though - it is a major step in the right direction. But, being able to buy browsing history, even in aggregate does not protect individuals.
- jt2190 10y ago>The zip code was intentionally not anonymized because it was seen as acceptable to release, and that is part of the issue of anonymized data. Forgive me for asking, but you seem to have two definitions of "anonymized": - anonymized - not anonymized, but claimed to be anonymized I think this argument (which I agree with) would be more forceful if we could stop calling non-anonymized data "anonymized". "Depersonalized", perhaps?
- satori99 10y agoAlso, the Netflix prize data set was used to identify individuals from only a handful of movie ratings. http://www.cs.cornell.edu/~shmat/shmat_oak08netflix.pdf http://www.cs.cornell.edu/~shmat/shmat_oak08netflix.pdf
- narrowrail 10y agoThe key part from that TD post for me was: [...]it requires a social media feed that includes a number of links to outside sites. However, they said that "given a history with 30 links originating from Twitter, we can deduce the corresponding Twitter profile more than 50 percent of the time." I'm paranoid enough to stay away from big social media altogether, but I realize that is uncommon.
- AdamSC1 10y agoWow only 30 links, that is both impressive and horrifying. I have very few tweets < 200, but I think you can find 30 outbound links in my first 50 posts since most of my use for it was sharing interesting articles. Can you imagine combining that with something else as simple as which Oauth apps someone has approved in Twitter? You'd reach near 100% accuracy in no time.