3 ms·
We had billions of Protobufs to store in Cassandra as byte blobs, using a zstd dictionary dramatically reduced storage size and improved latency over the built
by mbb70 3y ago
We had billions of Protobufs to store in Cassandra as byte blobs, using a zstd dictionary dramatically reduced storage size and improved latency over the built in compression. The complexity overhead of managing these dictionaries and making sure the client always has access to the right dictionary to decompress was non-trivial but well worth it.
We looked at Brotli as well but decompression speed at acceptable ratio was the most important factor for us, that plus the far superior docs and evangelism sealed the deal for zstd.
- metadat 3y agoWhat kind of additional gains (% wise) did you see with custom dictionaries compared to vanilla zstd?
- bhouston 3y agoDepends on how much shared entropy your data has. Could you test this by trying to compress all your content into one stream (shared dictionary) compared to compressing it into separate streams (no shared dictionary)?