3 ms·
Zstd's dictionaries contain two things: 1. Predefined statistics based on the training data for literals (bytes we couldn't find matches for), literal lengths,
by terrelln 7y ago
Zstd's dictionaries contain two things:
1. Predefined statistics based on the training data for literals (bytes we couldn't find matches for), literal lengths, match lengths, and offset codes. These allow us to use tuned statistics without the cost of putting the tables in the headers, which saves us 100-200 bytes.
2. Content. Unstructured excerpts from the training data that are very common. This gets "prefixed" the the data before compression and decompression, to seed the compressor with some common history.
Dictionaries are very powerful tools for small data, but they stop being effective once you get to 100KB or more.