4 ms·
HLL is great for a lot of reasons. He gives the primary reason as getting uniques from randomly sharded data in a distributed system. If your distributed syste
by brd529 10y ago
HLL is great for a lot of reasons. He gives the primary reason as getting uniques from randomly sharded data in a distributed system.
If your distributed system allows you to do a hash or range based sharding, for example by user_id, then you can do an accurate count(distinct user_id) across the system without a reshuffle of the data, knowing that all the data for a particular user lives on the same node.
- ozgune 10y ago(Ozgun from Citus Data) Yup, good point. This example shards the github_events table on user_id and then shows running count(distinct user_id). Since the sharding and count(distinct) column is the same, Citus can push down the count(distinct user_id) to each shard and then sum up the results from those shards. In this case, reshuffles don't come into the picture. In this example, HLLs would be most useful if the user then issued count(distinct repo_id) for example. Citus would then ask for HLL sketches from each shard and add them up on the coordinator node.