4 ms·
The math is correct, but I don't fully agree with the explanation. The saved bits don't really come from uniqueness, taking away 10,000 possibilities out of 2**
by codeflo 5y ago
The math is correct, but I don't fully agree with the explanation. The saved bits don't really come from uniqueness, taking away 10,000 possibilities out of 2**64 barely makes a dent. The combinatorial savings come almost entirely from ignoring order.
The real question is, what does a dataset have to have a unique list of identifiers be such a significat part of it that this is worth worrying about? Many identifier-heavy datasets have lots of cross references (e.g. n:m relation tables), with lots of repeated identifiers that are actually easy to compress. Am I missing something?