3 ms·
Really interesting. Curious about one particular case -- obfuscating the values of columns which are part of a `unique` constraint. Let's take, say, a set of
by rst 3y ago
Really interesting. Curious about one particular case -- obfuscating the values of columns which are part of a `unique` constraint.
Let's take, say, a set of a hundred thousand nine digit social security numbers, stored in a DB column that has a uniqueness constraint on it. This is a space small enough that hashing doesn't really mask anything -- there are few enough possible values that an adversary can compute hashes of all of them, and unmask hashed data. But the birthday paradox says that the RandomString transformer is highly unlikely to preserve uniqueness -- ask a genuinely random string generator to generate that many nine-digit strings, and you're extremely likely to get the same string out more than once, violating the uniqueness constraint.
One approach that I've seen to this is to assign replacement strings sequentially, in the order that the underlying data is seen -- that is, in effect, building up a dictionary of replacements over time. But that requires a transformer with mutable state, which looks kind of awkward in this framework. The best way I can see to arrange this is a `Cmd` transformer that holds the dictionary in memory. Is there a neater approach?
- ngalstyan4 3y agoNot sure what the approach of this library is, but can't you generate a nonce from a larger alphabet, hash the column values with the nonce `hash(nonce || column)`, and crypto-shred the nonce in the end. Then, during hashing you just need a constant immutable state, which effectively expands the hash space, without incurring the mutable state overhead of replacement strings strategy.