4 ms·
The x axis is compression speed, and the y axis is compression ratio. Zstandard outperforms zlib in compression ratio, compression speed, and decompression spe
by terrelln 8y ago
The x axis is compression speed, and the y axis is compression ratio.
Zstandard outperforms zlib in compression ratio, compression speed, and decompression speed (not shown). The only reason to stick with zlib is for compatibility with systems that expect zlib.
- jclay 8y agoThat sounds fantastic. What is the porting process generally like? Are there any possibilities to create an API compatible wrapper to make it a drop in replacement for zlib?
- Cyan4973 8y agoThere is a zlib wrapper included in the project : https://github.com/facebook/zstd/tree/master/zlibWrapper https://github.com/facebook/zstd/tree/master/zlibWrapper
- megous 8y agoRather simple. For example here's my port of qemu to use zstd compression algorithm for QCOW2 images, instead of zlib: https://megous.com/git/qemu-zstd/commit/?id=45df9c0510e737e6b2538401ec41f0f19f881b06 https://megous.com/git/qemu-zstd/commit/?id=45df9c0510e737e6...
- kstrauser 8y agoAwesome! What kind of metrics differences have you seen from the change?
- megous 8y agoMuch faster and/or better compression/decompression of images (tunable of course, depending on what level you set in the code), and also faster startup of VMs. I haven't measured exactly. I'd guess ~30% in my case. But all this depends on where the bottleneck is in any particular case. My VMs are on HDD, so increased decompression speed doesn't matter that much, but reduced size helps reading the necessary data faster. Linux VMs seem to be more compressible than Windows ones. It's certainly better than zlib in any case.
- terrelln 8y agoPorting is generally very easy. * If you already have the compression algorithm tagged, through a file extension, or a field, then you can use that to dispatch to the right decompression algorithm. * Zlib, gzip, xz, zstd, ... all have headers. If you are using zlib, and switching to zstd, you simply have to check the first 4 bytes for the zstd header using ZSTD_isFrame() [0], or attempting to decompress with zstd and if it fails fall back to the previous decompression algorithm. * The zstd CLI can decompress both zstd and zlib/gzip if compiled with zlib support. * Zstd provides a wrapper around the zlib API so you could transparently switch to zstd. [1] [0] https://github.com/facebook/zstd/blob/dev/lib/zstd.h#L1409 https://github.com/facebook/zstd/blob/dev/lib/zstd.h#L1409 [1] https://github.com/facebook/zstd/tree/dev/zlibWrapper https://github.com/facebook/zstd/tree/dev/zlibWrapper
- jclay 8y agoGreat, thanks! Last question. On the blog you mention you are underway porting the internal code to replace Zlib with Zstd. Is there a reason you decided not to use the wrapper as a first pass to migrate all uses of Zlib to Zstd across the entire codebase?
- terrelln 8y agoThere are a few reasons we haven't used the wrapper. * The larger services require tuning to get the best performance out of zstd, and we use some advanced options. * We have a "Managed Compression" library which does zstd dictionary compression, which doesn't work with the wrapper. * We have our own automatic decompression framework that handles many algorithms [0]. * A lot of use cases switched over from other algorithms than zlib. * A lot of use cases switched over to zstd organically, without our involvement, since it was such a clear win. [0] https://github.com/facebook/folly/blob/master/folly/compression/Compression.h#L522 https://github.com/facebook/folly/blob/master/folly/compress...
- hinkley 8y agoThe informatics dysfunction on this graph is just off the charts. Here's the problem. The graph is designed to make your conclusion sound right, but it doesn't actually prove that. Let's look at what the numbers really say: For the sample data, and the best case, gzip gets you a file that's about 31% of the original size. zstandard can get you a file that is 25% of the original size but it will take you four times as long to get it. If you allow it the same time as gzip, you can get 27% compression instead of 31%. That's only 13% improvement on the wire. That's nice, but it's not impressive at all. It's not a good enough reason to change your stack. The only thing that is impressive is that if you want the same compression ratio as gzip you can do it up to 20 times faster. On some hardware that's totally worth it, but not on all (because who is streaming at 700 MBps?)
- jclay 8y agoThere's so much that _feels_ misleading here. Paragraph before the chart: "The benefits we’ve found typically range from a 30 percent better ratio to 3x better speed." In what cases? What methodology was used to evaluate this? I certainly expected a scientific (honest) treatment of how the performance was evaluated and the corresponding trade-offs to follow on at some point.
- vjeux 8y agoIf you read the full article there are many examples of real systems being migrated with their associated wins.
- hinkley 8y agoI will say that we've been 'doing just fine' with zlib for longer than I'm comfortable with. I hung out on comp.compression when I was still in college and shared the dream of writing the next great compression library. Maybe the only thing I ever accomplished though was altering a Java minifier to improve compressibility of the class files. On a very deep level I'm pretty disappointed that zlib has been 'good enough' for almost 30 years. When it was Google proposing a change, I wasn't enthused about handing more control over HTTP to Google. They have too much already. The ways I'm concerned about Facebook have nothing to do with standards bodies or protocols. So maybe this is good enough. There are more options for files-in-motion and files-at-rest, and that could put it over the top. But you need to be filing high quality PRs on both HAProxy and Nginx if you want anybody to care.