4 ms·
Two things I'd like to say here 1. All anonymisation algorithms (k-anonymity, l-divergence, t-closeness, e-differential privacy, (e,d)-differential privacy, et
by float4 5y ago
Two things I'd like to say here
1. All anonymisation algorithms (k-anonymity, l-divergence, t-closeness, e-differential privacy, (e,d)-differential privacy, etc.) have, as you can see, at least one parameter that states to what degree the data has been anonymised. This parameter should not be kept secret, as it tells entities that are part of a dataset how well, and in what way, their privacy is being preserved. Take something like k-anonymity: the k tells you that every equivalence class in the dataset has a size >= k, i.e. for every entity in the dataset, there are at least k-1 other identical entities in the dataset. There are a lot of things wrong with k-anonymity, but at least it's transparent. Tech companies however just state in their Privacy Policies that "[they] care a lot about your privacy and will therefore anonymise your data", without specifying how they do that.
2. Sharing anonymised data with other organisations (this is called Privacy Preserving Data Publishing, or PPDP) is virtually always a bad idea if you care about privacy, because there is something called the privacy-utility tradeoff: you either have data with sufficient utility, or you have data with sufficient privacy preservation, but you can't have both. You either publish/share useless data, or you publish/share data that does not preserve privacy well. You can decide for yourself whether companies care more about privacy or utility.
Luckily, there's an alternative to PPDP: Privacy Preserving Data Mining (PPDM). With PPDM, data analysts can submit statistical queries (queries that only return aggregate information) to the owner of the original, non-anonymised dataset. The owner will run the queries, and return the result to the data analyst. Obviously, one can still infer the full dataset as long as they submit a sufficient number of specific queries (this is called a Reconstruction Attack). That's why a privacy mechanism is introduced, e.g. epsilon-differential privacy. With e-differential privacy, you essentially guarantee that no query result depends significantly on one specific entity. This makes reconstruction attacks impossible.
The problem with PPDM is that you can't sell your high-utility "anonymised" datasets, which sucks if you're a big boi data broker.
- amelius 5y agoCan advertisers be legally forced to use these mathematical techniques? Can we perhaps have a trusted third party which anonymizes data for these companies?
- cardosof 5y agoI don't think they could be legally forced to use specific algorithms - most laws state the ends ("privacy"), not the means. In the old world of analog marketing, you had market research companies (Nielsen, Kantar, GfK) to measure the audience and provide benchmarks. One way to help curb the power of adtech companies would be to force them to let go of measurement. That would require adjustments to privacy laws, creating a specific data processor for audiences role.
- int_19h 5y agoThe way it usually works, legislature writes the law that states the end, and establishes (or repurposes) an executive agency to implement the means, vesting them with the power necessary to do so. The agency then comes up with specific procedures etc - and they can enforce that. For example, the federal law in US does not define the procedure to properly destroy a firearm (such that it ceases to be regulated by the relevant laws) - but ATF does, and it's fairly specific: https://www.atf.gov/firearms/how-properly-destroy-firearms https://www.atf.gov/firearms/how-properly-destroy-firearms I don't see why the same approach couldn't work here.
- godelski 5y agoCould you not set legislation that to claim anonymity companies have to provide a certain level which is mathematically bounded? And/or as the OP suggested, add transparency? I'd think the combination is better since just the transparency won't make sense to most people and slow them to make informed decisions.
- b3morales 5y agoIn some ways it's preferable to leave the legislation a little open ended and put the details into a more flexible rule-making process. Then the rules can be updated by knowledgeable people as circumstances change, either to adopt new standards, relax them, or address unforeseen gaps.
- EGreg 5y agoI have a question Are zero-knowledge proofs really zero-knowledge if you do enough of them, can’t you reconstruct?
- cm2012 5y agoIt's kind of funny. Sure, with intense math you can maybe back into who some people are from an anonymous audience. Meanwhile the government just asks the ISP what you've been doing and they happily comply.
- llamataboot 5y agoThere is a very big difference here though, at least ostensibly (doesn't matter much to you if the government wants to know where you were two weeks ago at noon). The government has to prove based on the reasons that we choose in a democracy that we all want /why/ it has an interest in knowing such a thing. The companies, on the other hand, literally get to know whatever they want and it's up to us to prove why they shouldn't actually know that thing. Now, if we want to have a debate about which is more abused in practice, or which is more dangerous, I'm all about it. But the difference in access to information based on proving a need, versus proving a harm, is actually quite stark in theory.
- valenterry 5y agoDo you have a link where these concepts are explained in more detail?
- deleted 5y ago[deleted]
- motohagiography 5y agoImportant concepts. Key thing that has changed in privacy in last couple years is that de-identified data has recently been made into a legal concept instead of a technical one, whereby you do a re-identification risk assessment (not a very mature methodology in place yet), figure out who is accountable in the event of a breach, label the data as de-identified, and include the obligation of the recipients to protect it in the data sharing agreement. The effect on data sharing has been notable because nobody wants to hold risk, where previously "de-identification" schemes (and even encryption) made their risk and obligation evaporate as it magically transformed the data from sensitive to less sensitive using encryption or data masking. Privacy Preserving Data Publishing is sympathetic magic from a technical perspective, as it just obfuscates the data ownership/custodianship and accountability. FHE is the only candidate technology I am aware of that meets this need, and DBAs, whose jobs are to manage these issues, are notoriously insufficiently skilled to produce even a synthesized test data set from a data model, let alone implement privacy preserving query schemes like differential privacy. What I learned from working on the issue with institutions was nobody really cared about the data subjects, they cared about avoiding accountability, which seems natural, but only if you remove altruism and social responsibility altogether. You can't rely on managers to respect privacy as an abstract value or principle. Whether you have a technical or policy control is really at the crux of security vs. privacy, where as technologists we mostly have a cryptographic/information theoretic understanding of data and identification, but the privacy side is really about responsibilities around collection, use, disclosure, and retention. Privacy really is a legal concept, and you can kick the can down the road with security tools, but the reason someone wants to pay you for your privacy tool is that you are telling them you are taking on breach risk on their behalf by providing a tool. The people using privacy tools aren't using them because they preserve privacy, they use them because it's a magic feather that absolves them of responsibility. It's a different understanding of tools. However, it does imply a market opportunity for a crappy snakeoil freemium privacy product that says it implements the aformentioned techniques but barely does anything at all, and just allows organizations to say they are using it. Their problem isn't cryptographic, it just has to be sophisticated enough that non-technical managers can't be held accountable for reasoning about it, and they're using a tool so they are compliant. I wonder what the "whitebox cryptography" people are doing these days...