3 ms·
Slightly different domain, but I recently had to work on a system that retained email addresses that were known targets of aggressive spamming (450MM rows). In
by d2xdy2 11y ago
Slightly different domain, but I recently had to work on a system that retained email addresses that were known targets of aggressive spamming (450MM rows).
In a nutshell, the idea was to compare lists of emails and test to see how many accounts in the given list were also in the database. We'd typically get 10k item lists to test against the database.
Initially, I was creating a temporary table and doing a unique intersection on it. That took about 45 minutes on a modest machine. Next, I hacked up a bloom filter search (since all we cared about was whether or not the items in the sample were in the database); that ran in a little under a second for a set of 10k, but got a little unwieldy with sets approaching 1MM+.
We talked about it, and decided to just use Elasticsearch, and to just split the database into a bunch of buckets. Search time went up a small bit for small samples, but went waaay down for larger samples.
Also of note, going from spinning platters to SSD did make a pretty huge improvement in things when we were using a relational database. Now that it's all basically in memory, it doesn't matter so much.