4 ms·
The zip code was intentionally not anonymized because it was seen as acceptable to release, and that is part of the issue of anonymized data. US Zip codes serv
by AdamSC1 10y ago
The zip code was intentionally not anonymized because it was seen as acceptable to release, and that is part of the issue of anonymized data.
US Zip codes serve around 7,500 per code on average (US Census 2010) which is different than common wealth postal codes like in Canada where they serve an average of 19 households (25 - 75 people).
But, zip data can be interchanged with plenty of other unique identifiers on the web. Maybe it is browser language setting, or version of java etc.
Think of it like an Excel spread sheet, if in column A you can have options "1" or "2" then in a list of 100 people there will be at least 50 who share the same data footprint. If you keep adding columns from B onward with the same logic eventually you'll have pretty unique strings.
Things like searching history or web history are even worse. Ever done a search for a pizza place near your address, or Google map directions from your home to another location? That identifies you pretty easily. So does connecting to your works website, and the school your kid attends. Web browsing data is nearly impossible to anonymize by its nature unless it was compiled to something like "XX% of users in Zip XXXXXX visited website.com"
As for differential privacy, it is a nice emerging theory, but there are challenges with it as there is a significant trade off right now in terms of data accuracy when applying differential privacy. It is primarily effective at casting doubt on if variable "A" about user "B" in a data set is true or not, but if you don't have a specific target or specific metric then enough of the data is true that it could still in theory be deanonymized, and since most of the anonymity is based on incorrect variables in a data-set, all it would take to reverse engineer it is a large enough data-set and a few known variables.
I hope people like Apple continue to champion the advancement of differential privacy though - it is a major step in the right direction. But, being able to buy browsing history, even in aggregate does not protect individuals.
- jt2190 10y ago>The zip code was intentionally not anonymized because it was seen as acceptable to release, and that is part of the issue of anonymized data. Forgive me for asking, but you seem to have two definitions of "anonymized": - anonymized - not anonymized, but claimed to be anonymized I think this argument (which I agree with) would be more forceful if we could stop calling non-anonymized data "anonymized". "Depersonalized", perhaps?
- AdamSC1 10y agoSo there is the version of "anonymized" that most companies currently use which is as you say 'depersonalized' - they remove your personal information and think it is safe, but data is still tied together in unique records. (i.e. your name is replaced with an ID number) Then there is actually 'anonymized' data which would be the release of data in which you cannot in anyway identify a user. An example that comes to mind for me is the census releasing aggregate stats such as "14% of American's speak language X." If the census instead had records of each American line by line, listing which language they speak and other associated factors about them then this data is likely to paint a unique picture of the individual even if their personal information like name and address were removed. I think most data is very hard, if not impossible, to truly anonymize. Even if the search history that gets sold wasn't broken out into history per tuple/record, then you could still identify at least a few trends in it. Does that make more sense? But yes, I agree that these companies are more attempting to 'de-personalize' data for the sake of research, but, that is far from anonymous and naming it as such is misleading to the public.