9 ms·
Zstandard v1.4.0
- vbtechguy 7y agoPosted my zstd 1.4.0 compression benchmarks https://community.centminmod.com/threads/round-3-compression-comparison-benchmarks-zstd-vs-brotli-vs-pigz-vs-bzip2-vs-xz-etc.17259/ https://community.centminmod.com/threads/round-3-compression... - looking good :)
- jasonhansel 7y agoHow widely is this being adopted compared to (say) Brotli?
- stochastic_monk 7y agoI know it’s gained a lot of adoption in my circles, compared to none for Brotli. Several colleagues have written code using zstd as a library. Though anecdotal, certainly.
- terrelln 7y agoDisclaimer: I'm a maintainer of zstd, so I'm biased. Brotli dominates HTTP compression. Zstd just got its RFC approved a few months ago, but Brotli has been present in browsers for years. However, zstd is more widely adopted everywhere else, especially in lower level systems. Zstd is present in compressed file systems (BtrFS, SquashFS, and ZFS), Mercurial, databases, caches, tar, libarchive, package managers (rpm and soon pacman). There is a pretty complete list here https://facebook.github.io/zstd/ https://facebook.github.io/zstd/. Again, I'm biased because I know almost everywhere where zstd is deployed, but not everywhere that Brotli is.
- Svenstaro 7y agoWith the zstd support in pacman, arch will soon switch all packages over to zstd.
- lloeki 7y agoNitpick: it's actually in libarchive, pacman/makepkg doesn't really care about the format.
- jasonhansel 7y agoOut of curiosity: Why doesn't pacman just use HTTP's built-in compression? It could cache packages in gzipped form, but there's no reason to re-compress them over the wire.
- andrius4669 7y agoWhat would be gain?
- Svenstaro 7y agoBecause mirrors aren't guaranteed to use HTTP gzip compression and also gzip compresses much worse than xz.
- hackcasual 7y agoExcited for Zstd getting into browsers, Brotli is great for static assets, but it doesn't really outperform gzip by much at dynamic compression scenarios.
- JyrkiAlakuijala 7y agoIf you made your conclusion's from https://blog.cloudflare.com/results-experimenting-brotli/ https://blog.cloudflare.com/results-experimenting-brotli/ be warned that both of it's results and conclusions are flawed, because of a significant methodology flaw. You just cannot compare two things with a commutative function, and that blog post is based on an assumption that you can. Math just doesn't work like that. Brotli's fastest compression modes are 3-5x faster than zlib's fastest modes. For every gzip quality setting there is a brotli setting that is both faster and more dense than that gzip setting.
- hackcasual 7y agoWe used https://quixdb.github.io/squash-benchmark/ https://quixdb.github.io/squash-benchmark/ mostly. It's not that Brotli wasn't a winner compared to our gzip, it just didn't outperform it substantially and came with challenges integrating it with our product for supporting dynamic payloads. Meanwhile static assets just need brotli support at build time.
- JyrkiAlakuijala 7y agoWhile it is a great benchmark, that is about 3 year old data. The brotli quality 1 was moved to be quality 2, and two new levels have been added. Everything has been sped up and a few levels (possibly 5-10) have got a ~5 % density boost from improved context modeling. Level 10 has been added (about 2-3x faster than level 11) -- in squash benchmark you still see the initial behavior where level 10 copied level 11 performance.
- JyrkiAlakuijala 7y agoI'm also biased, as I'm the author of Brotli. With Brotli you get about 5 % more density. Brotli decompresses about 500 MB/s while Zstd decompresses about 700 MB/s. Typical web page is 100 kB, so you need to wait 200 us for decompression. For zstd, you'd only need to wait 140 us. That 60 us saving comes with a cost: you'll be transferring 5 % more bytes, which can cost you a hundreds of milliseconds. Brotli is also more streaming, so you get your bytes our earlier during the data transfer. This allows for pages to be rendered with partial content and further fetches to be issued earlier. Zstd supporters have used comparisons against brotli where they compare a small window brotli against large window zstd. This makes it seem like zstd can compete in density, too, but that is just apples to oranges comparisons.
- terrelln 7y agoBrotli is great at compressing static web content, zstd without a dictionary is unlikely to outperform it. For static content you'd probably rather save 5% of space over some decompression costs, since Brotli decompression is fast enough. Zstd has an advantage if you don't have the CPU to compress at the maximum level, since zstd is generally faster than Brotli at the lower levels. Even still, for web compression, Brotli has the advantage of already being present in the browsers, so you're betting off using Brotli for web compression as it stands today.
- JyrkiAlakuijala 7y agoFor encoding there shouldn't be a format specific difference. If there is, it is based on immaturity of the implementations. There were stages when brotli:0 used to be faster to compress than any setting in zstd, but now they have played catch-up and are the leader. Zstd can be significantly slower in medium levels -- you just need not to be tricked to apply compressors at different window sizes. Zstd changes window sizes under the hood. With brotli you need to explicitly decide about your decoding resource use. If you use the same window size and aim for the same density of compression, brotli actually tends to compress faster in the middle qualities, too.
- ryacko 7y agoHTTP compression should be disabled by default, 99% of web content downloaded is already compressed.
- algorithmsRcool 7y agoBut this is flatly untrue. Images and video sure, but everything from html, just, css, svg isn't compressed. In fact compression is critical to modern SPA frameworks to keep initial download times lower. Also this post has little to do with http compression. ZSTD is used it many other circumstances.
- deathanatos 7y agoAnd for resources that don't compress well, web servers like nginx (and I presume others) support listing what mimetypes to compress, so they won't double-compress those things.
- znedw 7y agoI've been using zstd compression on btrfs for a while now and it's excellent, most of my stuff is already compressed (movies, music) but my home directory (which is mostly comprised of text files) has shrunk greatly.
- terrelln 7y agoThe next GRUB release (grub-2.04) includes my patch to add support for zstd compressed BtrFS filesystems, which should solve one of the major pain points of Zstd BtrFS compression.
- znedw 7y agoI tried manually building grub with the patch but gave up and just created an ext2 /boot partition, but this is good news! Thank-you for fixing it.
- cmurf 7y agoCool! Looking forward to that. A temporary work around in the meantime, I've used `chattr +C` on the directories I want to be exempt from zstd compression, so that grub can read those files.
- _wmd 7y agoZstandard is awesome, there is nothing else close to it that I've seen. There are plenty of codecs that give you either obscene ratios or low CPU usage, but none that support both in combination. Zstandard on its highest setting easily competes with xzip for anything I've thrown at it. Decompression throughput is unaffected by compression setting -- it just burns more memory. So you can get xzip-like ratios with decompression throughput approaching 1GB/sec It allows trading lower CPU for gzip-or-worse compression, and you can mix the settings within a single file. This means you can e.g. use the lowest setting (or no compression at all - it supports that too) to append to a file, while occasionally recompressing recent appends into a single block using the highest setting -- so the cost of compression can be amortized The only petty annoyance with it is ecosystem support - e.g. GNU tar has no option for it, so it's slightly more painful to work with
- jacobolus 7y ago> nothing else close to it that I've seen My understanding is that Oodle generally performs better (less CPU on both ends for any desired compression ratio) than ZStd in pretty much every context, http://www.radgametools.com/oodle.htm http://www.radgametools.com/oodle.htm ... assuming you control both compression and decompression and are willing to spend money for it.
- pcwalton 7y agoI don't see comparisons against any version of Zstandard, much less Zstandard 1.4, there. (Besides, regardless of the state of benchmarks today, the momentum is clearly with Zstandard, with Facebook, Intel, and the open source community behind it. The basic lossless compression algorithms haven't changed much since the late 1970s. Making compression fast is mostly just a long slog of engineering hurdles, the kind that big companies are very good at doing.)
- yoklov 7y agohttp://cbloomrants.blogspot.com/2018/06/zstd-is-faster-than-leviathan.html http://cbloomrants.blogspot.com/2018/06/zstd-is-faster-than-... is versus zstd 1.3.3. There are others too on that blog. Oodle and zstd are both very impressive, it's a shame the former is not free.
- ac29 7y agoAnyone know a decent Windows implementation? There's a 7zip fork that includes Zstd support, but it can only put Zstd inside a .7z container, which doesn't appear to work with any other tools.
- svnpenn 7y agoThe linked site has windows builds...
- ac29 7y agoTrue! I suppose what I actually want is a Windows utility that can make a .tar.zst archive, ideally from a GUI. In the Windows world, archiving and compression are usually in a single file type (.zip, .rar, .7z). Zstd follows the unix style where it can't directly compress folders of files, they need to be in a tar (or other archive format) first. This isn't really an issue on Linux, since Zstd support is built into tar, which ships on pretty much every system.
- JyrkiAlakuijala 7y agoThere are two ways of doing compression or archives. One where a local small corruption destroys one data entity, and another where a local small corruption destroys most if not all the archive. The latter kind gives a small density improvement, but can prove to be the wrong option some time later.
- chungy 7y agoThere haven't yet been any extensions to zip or 7z for zstd support. There is a branch of wimlib that has experimental zstd support, though it's unlikely it will ever be merged into the master branch. You could make uncompressed zip or 7z files and compress that independently as a zst file, but that's a bit baroque compared to just using tar. :) 7-Zip does often seem bent on supporting everything, I imagine some day in the future it'll support zstd at least as an independent archive, if not extending the Zip and 7z formats as well.
- albertzeyer 7y agoOlder discussion: https://news.ycombinator.com/item?id=16226923 https://news.ycombinator.com/item?id=16226923 Some of the obvious competitors are brotli and snappy. Here some comparison: https://quixdb.github.io/squash-benchmark/ https://quixdb.github.io/squash-benchmark/ http://www.mattmahoney.net/dc/text.html http://www.mattmahoney.net/dc/text.html (I wonder if anyone has an updated picture with Pareto front for these numbers.)
- JyrkiAlakuijala 7y agoMahoney's benchmark is missing the large-window brotli numbers, which are about 5 % better than those of zstd and 10 % better than those of small-window brotli. Brotli with large window can do around 199M for the 1G text corpus: https://groups.google.com/forum/m/#!topic/brotli/aq9f-x_fSY4 https://groups.google.com/forum/m/#!topic/brotli/aq9f-x_fSY4 Here a large window aggregate result view: https://encode.ru/threads/2947-large-window-brotli-results-available-for-a-several-corpora https://encode.ru/threads/2947-large-window-brotli-results-a...
- drej 7y agoThere are basically three scenarios where I choose various compression algorithms (a few exceptions excluded): - maximum compatibility (while tolerating low performance) - gzip - great performance (while tolerating larger files) - snappy - very good performance with good (not best) compression ratios - zstd I don't really want to use Any New Shiny Algo to compress some data that might outlive this piece of software, that's why I use gzip very often, because I know I'll always be able to decompress it. But I've been increasingly adopting zstd and snappy for one single reason - they are becoming widely supported within the ecosystems I work in (data processing). That, to me, is more important than compression ratios and decompression speeds.
- d33 7y agoHow does snappy compare to lz4? Also, for gzip, you might want to consider using its multithreaded version, called "pigz".
- terrelln 7y agolz4 compresses and decompresses faster than snappy, and compresses similarly. You can see some comparisons on the GitHub's readme https://github.com/lz4/lz4 https://github.com/lz4/lz4.
- coldpie 7y agoStrangely, snappy doesn't seem to ship a binary on Arch Linux, just a shared library. Is there no snappy command-line program?
- usefulcat 7y agoWe have many terabytes of large files that are currently compressed using xz (well, pixz actually). In terms of compression speed and ratio, zstd is pretty comparable, and for single-threaded decompression it's faster. At this point the only thing stopping us from using it instead of xz/pixz is the fact that multi-threaded pixz decompression is faster. Are there any plans to add MT decompression to zstd?
- cmurf 7y agoOn three test machines, HDD and SSD, the single decompression thread ranges from 9% to 15% CPU, and has maxed out the read+write capacity for my storage. But maybe you have super fast source and target storage?
- usefulcat 7y agoIn my scenario, I'm decompressing from HDD and piping the decompressed data directly to another process. I may be limited by disk sometimes, but often the data is already in the filesystem cache. To give a concrete example: zstd -d < compressed.zst | pv > /dev/null == ~330 MB/s For comparison, pixz with the same data using 32 cores: pixz -d -p 32 < compressed.xz | pv > /dev/null == 1.15 GB/s Granted, zstd is far, far more efficient per core, but there are plenty of workloads where I can afford to use a lot of cores for decompression. Also pixz still compresses slightly better than zstd -19, but I'd be willing to trade that for more efficient decompression if I could still have the option of really fast decompression using multiple threads. Note also that with this particular data, I'm seeing a compression ratio of only about 4.3:1 using zstd -19. I can imagine that zstd would use less CPU when decompressing if the ratio was higher.
- cmurf 7y agopzstd does support -d (decompression) option with a default of 4 processes, which can be set with -p. It's part of the zstd package, but I guess it's separate because it's experimental? Not sure what the difference is between zstd -T4 for compressing, and pzstd -p4 for compressing. Anyway, for a test file I get ~125% CPU with pzstd -d, so it is able to do more work, and slightly decreases time. Decompress zstd -d real 0m19.125s user 0m10.888s sys 0m2.575s Decompress pzstd -d real 0m14.819s user 0m14.017s sys 0m4.163s