3 ms·
Am I correct in thinking the accuracy of this method has more to do with the distribution in the sample than the algo itself. Meaning, given a 100% unique list
by deadeye 2y ago
Am I correct in thinking the accuracy of this method has more to do with the distribution in the sample than the algo itself. Meaning, given a 100% unique list of items, how far off would this estimate be?
- MarkusQ 2y agoIt doesn't appear to be heavily dependent on the distribution. It mostly depends on how large your buffer is compared to the number of unique items in the list. If the buffer is the same size or larger, it would be exact. If it's half the size needed to hold all the unique items, you drop O(1 bit) of precision; at one quarter, O(2 bits), etc.
- edflsafoiewq 2y agoI don't think so. I think the point of the removal in X <- X \ {a_i} With probability p, X <- X U {a_i} is that afterwards a_i is in X with probability p regardless of how many times it occurs in the stream.