5 ms·
I must be missing something. You have 2MB of data for "is my 6 character SHA1 prefix in any breach". Why can't you distribute that to every edge nodes ?
by smallpipe 1y ago
I must be missing something. You have 2MB of data for "is my 6 character SHA1 prefix in any breach". Why can't you distribute that to every edge nodes ?
- lerp-io 1y agocan’t u just store single hash and use bloom filter or something to check if ur email is in hash on the client side also (or maybe that’s what they are doing and don’t wanna send the large data if it’s several mb idk)
- lerp-io 1y agoi just checked and ai said bloom filter is faster and more efficient than k-anon lookup, maybe in the next article lol.
- Thorrez 1y agoThere are tons of emails that share the same prefix. When you lookup a prefix, you can't simply get a boolean response. You have to get a list of emails as the response. The client then searches through the list to see if the desired email is in the list or not. Returning a list of emails instead of a single bit significantly increases the data size. Additionally, people don't just want a boolean answer of "was my email breached somewhere". They want a list of all the breaches that breached the email. So the returned data actually needs to be a list of emails and the list of breaches that each email was breached in. >Via the public API. This endpoint also takes an email address as input and then returns all breaches it appears in.
- qw 1y ago> The client then searches through the list to see if the desired email is in the list or not. The initial prefix check would probably reduce the amount of lookups necessary, as it would only be necessary to do a deeper search if the prefix matches.
- smallpipe 1y agoYeah that was my point, you can get rid of a significant portion of requests at the edge with a bloom filter, and there's no reason you have to build the bloom filter locally as requests come in. Instead, it can be created ahead of time, when the dataset is updated.
- Thorrez 1y agoSee my reply at https://news.ycombinator.com/item?id=43780713 https://news.ycombinator.com/item?id=43780713 . Also regarding "you can get rid of a significant portion of requests at the edge with a bloom filter", Troy's existing design already gets rid of a significant portion of requests at the edge. That's why he says >The response from each search was coming back so quickly that the user wasn’t sure if it was legitimately checking subsequent addresses they entered or if there was a glitch.
- Thorrez 1y ago>only be necessary to do a deeper search if the prefix matches There are 5 billion emails in at least 1 breach and 16 million prefixes. Almost all if not all prefixes have at least 1 email in a breach. So almost all prefixes match. I don't see why it's useful to spend a bunch of effort optimizing the very rare case of a prefix not matching. Now, if the bloom filter checked emails instead of checking prefixes, that would be useful. However, a bloom filter of 5B elements with a 10% false positive rate would be 2.8 GB, which is prohibitively large. https://hur.st/bloomfilter/?n=5g&p=10&m=&k= https://hur.st/bloomfilter/?n=5g&p=10&m=&k=
- Thorrez 1y agoWhere did you get the number 2MB? According to this calculator, a bloom filter for 16M elements with a 10% false positive rate would be 9MB. https://hur.st/bloomfilter/?n=16m&p=10&m=&k= https://hur.st/bloomfilter/?n=16m&p=10&m=&k=