4 ms·
> Not sure what the state the art is in searchable encryption for db indexes, but just trying to do stuff that requires a scan becomes untenable due to having t
by CiPHPerCoder 2y ago
> Not sure what the state the art is in searchable encryption for db indexes, but just trying to do stuff that requires a scan becomes untenable due to having to read and decrypt on the client to find it or aggregate it.
There are a lot of different approaches, but the one CipherSweet uses is actually simple.
First, take the HMAC() of the plaintext (or of some pre-determined transformation of the plaintext), with a static key.
Now, throw away most of it, except a few bits. Store those.
Later, when you want to query your database, perform the same operation on your query.
One of two things will happen:
1. Despite most of the bits being discarded, you will find your plaintext.
2. With overwhelming probability, you will also find some false positives. This will be significantly less than a full table scan (O(log N) vs O(N)). Your library needs to filter those out.
This simple abstraction gives you k-anonymity. The only difficulty is, you need to know how many bits to keep. This is not trivial and requires knowing the shape of your data.
https://ciphersweet.paragonie.com/security#blind-index-information-leaks https://ciphersweet.paragonie.com/security#blind-index-infor...
I proposed the same technique to AWS, who adopted it under the name Beacons for the AWS Database Encryption SDK.
https://docs.aws.amazon.com/database-encryption-sdk/latest/devguide/beacons.html https://docs.aws.amazon.com/database-encryption-sdk/latest/d...
- adunsulag 2y agoI was reading your reply and started thinking, this sounds a lot like what I did to do encrypted search with Bloom Filters and indexes. I click on the first link and find the exact website I used when researching and building our encrypted search implementation for a health care startup. It worked fabulously well, but it definitely requires a huge amount of insight into your data (and fine-tuning if your data scales larger than your initial assumptions). That's awesome that AWS has now rolled it into their SDK. I had to custom build it for our Node.JS implementation running w/ AWS's KMS infrastructure. Are you the author of the paragonie website? The coincidence was startling. If so, I greatly thank you for the resource. Edit After going back and re-reading the blog post, looks like you are the author. Again thank you, you were super helpful .
- CiPHPerCoder 2y ago> Are you the author of the paragonie website? The coincidence was startling. If so, I greatly thank you for the resource. Thanks. Yes, I'm one of the authors.