3 ms·
In short, we work with individual teams and projects to evaluate their priorities, and select or build the compression scheme that best addresses their needs. S
by felixhandte 8y ago
In short, we work with individual teams and projects to evaluate their priorities, and select or build the compression scheme that best addresses their needs. Some of those cases are described in the post, but we wanted to summarize the general kinds of results we see.
I did spend a fair amount of time working on a graphic for the post to try to capture the distribution of improvements we've seen across different use cases at Facebook, but it ended up being very hard to interpret / glean anything meaningful from.
It's difficult to present rigorous conclusions: the reality of data compression is that every use case is different, with different priorities and data characteristics that cause different compressors to behave differently.
Some data (random noise) is totally incompressible, and zstd will do just as poorly as any other algorithm. On the other hand, with highly structured / repetitive data like JSON, zstd can do enormously better than zlib. It all depends on the specific context.
Another example of how benchmarking is hard: we recently spent a fair amount of time improving zstd for highly-contended memory scenarios [1]. This work basically won't show up in a standard single-threaded single-workload benchmark. But we saw meaningful improvements in the real world at Facebook.
Ultimately, if you are evaluating using zstd (or any other compressor), the best predictor for performance will be a benchmark you run yourself with your own data. And hopefully you like what you see!
[1] https://github.com/facebook/zstd/releases/tag/v1.3.6 https://github.com/facebook/zstd/releases/tag/v1.3.6