3 ms·
The technical details vary a lot, but the basic idea that data can be meaningfully separated is extremely common belief. Most people - even many people with tec
by pdkl95 5y ago
The technical details vary a lot, but the basic idea that data can be meaningfully separated is extremely common belief. Most people - even many people with technical backgrounds - don't have personal experience using powerful modern tools like SQL JOIN. I like DJB's description[1] of the problem:
>> Hashing is magic crypto pixie-dust, which takes personally identifiable information and makes it incomprehensible to the marketing department. When a marketing person looks at random letters and numbers they have no idea what it means. They can't imagine that anybody could possibly understand the information, reverse the hash, correlate the hashes, track them, save them, record them.
In general this isn't an intentional deception. People probably legitimately believe they are protecting data by "anonymizing" (hashing) certain fields. The hashes in the "anonymized" data look impossibly large. The fact that modern computers can simply enumerate all possible hashes in a few minutes[2] is not immediately obvious.
> the systems are black boxes
The need for algorithmic transparency will only increase over time. Unfortunately, far too many corporations rely on being able to launder private data as their business model.
[1] https://projectbullrun.org/surveillance/2015/video-2015.html#bernstein https://projectbullrun.org/surveillance/2015/video-2015.html...
[2] https://www.theguardian.com/technology/2014/jun/27/new-york-taxi-details-anonymised-data-researchers-warn https://www.theguardian.com/technology/2014/jun/27/new-york-...
- ssss11 5y agoComing from a technical background I’m still dumbfounded that this exists. I actually think that the need for algorithmic transparency will result in some industry body that audits closed source code, db schemas and infra/architecture and provides a public rating system.. I’d love to see that. Edit: typo