5 ms·
Zstandard is awesome, there is nothing else close to it that I've seen. There are plenty of codecs that give you either obscene ratios or low CPU usage, but non
by _wmd 7y ago
Zstandard is awesome, there is nothing else close to it that I've seen. There are plenty of codecs that give you either obscene ratios or low CPU usage, but none that support both in combination. Zstandard on its highest setting easily competes with xzip for anything I've thrown at it. Decompression throughput is unaffected by compression setting -- it just burns more memory. So you can get xzip-like ratios with decompression throughput approaching 1GB/sec
It allows trading lower CPU for gzip-or-worse compression, and you can mix the settings within a single file. This means you can e.g. use the lowest setting (or no compression at all - it supports that too) to append to a file, while occasionally recompressing recent appends into a single block using the highest setting -- so the cost of compression can be amortized
The only petty annoyance with it is ecosystem support - e.g. GNU tar has no option for it, so it's slightly more painful to work with
- jacobolus 7y ago> nothing else close to it that I've seen My understanding is that Oodle generally performs better (less CPU on both ends for any desired compression ratio) than ZStd in pretty much every context, http://www.radgametools.com/oodle.htm http://www.radgametools.com/oodle.htm ... assuming you control both compression and decompression and are willing to spend money for it.
- pcwalton 7y agoI don't see comparisons against any version of Zstandard, much less Zstandard 1.4, there. (Besides, regardless of the state of benchmarks today, the momentum is clearly with Zstandard, with Facebook, Intel, and the open source community behind it. The basic lossless compression algorithms haven't changed much since the late 1970s. Making compression fast is mostly just a long slog of engineering hurdles, the kind that big companies are very good at doing.)
- yoklov 7y agohttp://cbloomrants.blogspot.com/2018/06/zstd-is-faster-than-leviathan.html http://cbloomrants.blogspot.com/2018/06/zstd-is-faster-than-... is versus zstd 1.3.3. There are others too on that blog. Oodle and zstd are both very impressive, it's a shame the former is not free.
- _wmd 7y agoZstandard uses ANS in some (all?) of its modes, that's very much brand new -- 2014
- jacobolus 7y ago> regardless of the state of benchmarks today, the momentum is clearly with Zstandard, with Facebook, Intel, and the open source community behind it. This argument by hand-wave doesn’t match up with demonstrated progress over the past few years. Irrespective of the skill or insight of its individual engineers, I would be surprised if a schizophrenic hack-it-with-duct-tape kind of engineering culture like Facebook could keep up with a focused and motivated expert like cbloom over the medium term (though I suppose the latter could conceivably at some point lose interest in the domain and switch to building something else).
- pcwalton 7y agoI've seen this play out before. The history of jemalloc, particularly how it rapidly outpaced virtually every other allocator while under development at Facebook, suggests otherwise. At this point jemalloc is so good that there is little reason other than NIH to use anything else (unless you want an especially hardened allocator for security and are willing to give up some performance). Even Google uses it in Android. Zstandard is likewise deservedly on track to dominate the lossless compression space.
- vthriller 7y ago> At this point jemalloc is so good Some folks disagree: > [Ruby's] memory usage only reduces when using jemalloc 3; memory usage is still high when using jemalloc 5. Nobody knows why, so that makes the choice of defaulting to jemalloc very dodgy. via https://www.joyfulbikeshedding.com/blog/2019-03-29-the-status-of-ruby-memory-trimming-and-how-you-can-help-with-testing.html https://www.joyfulbikeshedding.com/blog/2019-03-29-the-statu...
- mappu 7y agoZstd is from Facebook, sure, but to be more specific its development is lead by Yann Collet (of LZ4 fame) who is unquestionably a focused and motivated expert.
- fuzzy2 7y agoBut it’s not WinZip or WinRAR. This is a product marketed for professional use(r)s. How could it ever become more popular than something you just have?
- terrelln 7y agotar-1.3.1 added support for zstd with the option `--zstd` and `-a`, the auto-decompression flag, also supports zstd. Older tar versions also have the `-I` flag which you can use to (de)compress with zstd. We're working to improve zstd support in the ecosystem over time, but this work moves slowly, and it takes a long time for upstream work to make it to users systems, especially LTS systems.
- feanaro 7y ago> tar-1.3.1 added support for zstd with the option `--zstd` [...] Is there a way to pass options to zstd when used like this?
- aepiepaey 7y agoLast time I checked zstd does not support getting options from an environment variable (unlike .e.g. gzip). With a new enough version of tar you can pass a command with arguments to '-I' (e.g. `tar -I 'zstd -19' ...`). The alternative is piping the output from tar through zstd yourself.
- vbtechguy 7y agoI build my own custom tar 1.32 rpm with zstd support https://community.centminmod.com/threads/custom-tar-archiver-rpm-build-with-facebook-zstd-compression-support.16243/ https://community.centminmod.com/threads/custom-tar-archiver... and can confirm you can pass options that way.
- terrelln 7y agoStarting with zstd-1.3.8 we support the `ZSTD_CLEVEL` environment variable. We've started with a small scope for the variable, because we don't want users to unexpectedly remove the source fie, for instance. If you want to pass extra options, you can pipe the output to zstd, which is exactly what tar is doing internally.
- Twirrim 7y agoI got to experiment with it on some pretty poor spec MIPS processors not long after it was made public. Even there, on architecture it wasn't designed of specifically optimised for, it seriously outperformed competitors.
- therealmarv 7y agoCould be that Zstandard outperformes it but I also have similar good experience with brotli. Good compression speed ratio (use q 1 for brotli lower than 0.6): tar -I"brotli -q 2" -cvf file.tar.br inputfiles Decompression: tar -Ibrotli -xvf file.tar.br It outperformes gzip and is near xz region without a lot of CPU power and is also blazing fast. Really useful if you have e.g. a big PostgreSQL db dump which you want to transfer to your own machine. Examples for Postgres dumping and restoring: pg_dump adatabase | brotli -q 2 > dump.sql.br brotli -d < dump.sql.br | psql -U postgres dbusernamehere
- vthriller 7y ago> There are plenty of codecs that give you either obscene ratios or low CPU usage, but none that support both in combination. I still find lbzip2 (which is a bzip2 reimplementation with better algorithms and support for multithreading) quite competitive for highly compressible data. Here's quick and unscientific test that shows that lbzip2 (-9) is still both faster and has better ratio than zstd (-12, --long or not) while also using the least amount of RAM (tmpfs, multithreaded compression using 4-core Xeon E5-2603 v1): $ time lbzip2 -k linux-5.0.8.tar real 24.21 user 84.87 sys 4.07 maxrss 105472 $ time ~/zstd-1.4.0/zstd -T0 -k -12 linux-5.0.8.tar -o linux-5.0.8.tar.zst-12 real 30.69 user 105.27 sys 0.61 maxrss 942416 $ time ~/zstd-1.4.0/zstd -T0 -k -12 --long linux-5.0.8.tar -o linux-5.0.8.tar.zst-12-long real 31.28 user 107.90 sys 0.86 maxrss 1532432 $ time xz -T0 -k -2 linux-5.0.8.tar real 34.40 user 123.59 sys 0.57 maxrss 410192 $ stat -c '%s %n' linux-5.0.8.tar* | sort -n 126382954 linux-5.0.8.tar.bz2 126394210 linux-5.0.8.tar.zst-12-long 128003669 linux-5.0.8.tar.zst-12 131418488 linux-5.0.8.tar.xz 863426560 linux-5.0.8.tar The only clear advantage zstd has is decompression speed: $ time xzcat -T0 linux-5.0.8.tar.xz >/dev/null real 17.25 user 17.06 sys 0.17 maxrss 17312 $ time ~/zstd-1.4.0/zstd -dc -T0 linux-5.0.8.tar.zst-12 >/dev/null real 2.08 user 1.97 sys 0.08 maxrss 27088 $ time ~/zstd-1.4.0/zstd -dc -T0 linux-5.0.8.tar.zst-12-long >/dev/null real 2.26 user 2.03 sys 0.17 maxrss 535360 $ time lbzcat linux-5.0.8.tar.bz2 >/dev/null real 10.34 user 33.74 sys 3.53 maxrss 127088
- aldanor 7y agoGiven that it's the decompression speed that is typically the "user-facing" part in many contexts (with compression being done by automatic jobs etc), such a difference in decompression speed is pretty awesome indeed. (which is also what makes it a near-perfect codec for HDF5, via blosc-hdf5)
- dancek 7y agoI mostly use lbzip2 for day-to-day tasks too. But shouldn't you compare to pzstd to be fair? It's not very surprising that running 4 threads is faster than 1 thread in wall clock time. EDIT: no, I was wrong, `zstd -T0` is basically the same as `pzstd`.
- cout 7y agoIn my experience, for the best balance between compression speed and compression ratio, nothing beats 7zip with the right options: -mmt=$(nproc) # use all available cores -ms=off # disable solid archives (compress each file separately) -m0=lzma2 # lzma2 has better threading than lzma1 -md=64m # dictionary size -ma=0 # "fast" mode -mmf=hc4 # hash chain match finder -mfb=64 # number of "fast bits" -mf=off # disable filters The biggest gains are: 1) using all available cores, 2) setting the match finder (the binary tree match finders are terribly slow; I haven't played much with the newer patricia tree match finders), 3) disabling solid archives (this seems to cause 7zip to distribute the work more evenly between cores, though it still may only use a few cores if there are many small files), 4) using "fast" mode (whatever that is, it gives a noticeable performance boost and doesn't seem to affect compression ratio much). Every few years I try zstd and others, and for the data I work with (primarily a mix of json and fixed-width-field binary data), lots of tools beat 7zip out of the box, but they fall short of 7zip with the above command-line options.
- cmurf 7y agoPerhaps it's the use of a dictionary? As far as I'm aware, tar, zstd, xz do not use one by default, it's an extra set of hoops to create a training set, create the dictionary, use it for compression, and pack it away somewhere so that it's available for decompression, and then actually use it for decompression. If that's all being done by 7zip just by passing -md=64m that's pretty cool. Edit: Ahh, I was confused. Neither require a separate training step. Zstd offers an option to do a training step. Both always use dictionaries with a default size that can optionally be changed.
- terrelln 7y agoA comparable zstd call that uses a 64 MB window size and all cores is: zstd --long=26 -T0 From there you can tune the compression level, or increase the window size up to 2 GB (--long=31). zstd won't beat the compression of xz, but it can compress much faster if you trade off some space.
- vthriller 7y agoWell, it didn't work that well for me. Adding to https://news.ycombinator.com/item?id=19682019 https://news.ycombinator.com/item?id=19682019 $ time 7zr a -mmt=$(nproc) -ms=off -m0=lzma2 -md=64m -ma=0 -mmf=hc4 -mfb=64 -mf=off linux-5.0.8.tar{.7z,} real 60.49 user 158.94 sys 3.06 maxrss 8995040 $ stat -c '%s %n' linux-5.0.8.tar.7z 127700475 linux-5.0.8.tar.7z $ time 7zr e -so linux-5.0.8.tar.7z >/dev/null real 14.09 user 13.96 sys 0.12 maxrss 282208 Basically: - it took twice the time to compress data even compared to xz -2 (which also uses lzma2 under the hood), - it is comparable to zstd/bzip2 ratio-wise, - it used almost 6 times (!) more RAM than even zstd -12 --long, - it only used about 2.5 CPU cores out of 4 while compressing (which aligns pretty well with your reasoning for using -ms=off). ---- But hey, source code is not that regular. Since you mentioned JSON and fixed-width-field binary data, I decided to re-run benchmarks on 10M lines of nginx access logs: they're way more regular in their structure (repetitive URLs, timestamps, Mozilla/5.0, stuff like that) that might benefit from larger window sizes. $ time lbzip2 -k access-log-10m.log real 90.59 user 313.04 sys 18.46 maxrss 117904 $ time ~/zstd-1.4.0/zstd -T0 -k -12 access-log-10m.log -o access-log-10m.log.zst-12 real 77.34 user 277.21 sys 1.55 maxrss 886416 $ time ~/zstd-1.4.0/zstd -T0 -k -12 --long access-log-10m.log -o access-log-10m.log.zst-12-long real 69.24 user 242.18 sys 1.85 maxrss 1411872 $ time 7zr a -mmt=$(nproc) -ms=off -m0=lzma2 -md=64m -ma=0 -mmf=hc4 -mfb=64 -mf=off access-log-10m.log{.7z,} real 109.10 user 356.42 sys 4.69 maxrss 9777664 $ stat -c '%s %n' access-log-10m.log* | sort -n 208537395 access-log-10m.log.bz2 231953002 access-log-10m.log.zst-12-long 237566691 access-log-10m.log.zst-12 249412192 access-log-10m.log.7z 3386733539 access-log-10m.log Now tweaked 7z did better CPU- and time-wise, but it's still behind zst and bz2 on every metric, especially RAM which it requires so much of (literally gigabytes) it becomes impractical in a number of situations. And we needed a pretty regular input (not just some pretty compressible text like source code or Wikipedia dump) to close that gap. So I can't really recommend your suggestion, unless you have some niche input that benefits from that particular set of options (but then, who has time to learn lzma internals and how every option plays with different kinds of input?). ---- It's also worth pointing out that zstd also has plenty of options to fiddle with: https://github.com/facebook/zstd/blob/dev/programs/zstd.1.md#advanced-compression-options https://github.com/facebook/zstd/blob/dev/programs/zstd.1.md...