8 ms·
A parallel implementation of gzip for modern multi-processor multi-core machines
- spullara 10y agoMy guess is that this compresses less efficiently as you would have to shard the dictionaries. Might be close though for large files. I was surprised that there were no speed or efficiency comparisons in the README.
- kbaker 10y agoThe max window size for zlib is 32 KB, so I don't think the default sharding at 128 KB would change much. You can pass the -b parameter if you find out a bigger shard works better on your data. If you are looking for details of the design of pigz, there is a very well-documented overview in the source of pigz.c: https://github.com/madler/pigz/blob/master/pigz.c#L187 https://github.com/madler/pigz/blob/master/pigz.c#L187
- chii 10y agohad a quick scroll thru that source code - i didnt know you could implement a try/catch in C via macros...mind blown.
- TimJYoung 10y agoThis book (C Interfaces and Implementations: Techniques for Creating Reusable Software) has a whole section on it: https://www.amazon.com/Interfaces-Implementations-Techniques-Creating-Reusable/dp/0201498413 https://www.amazon.com/Interfaces-Implementations-Techniques...
- hrehhf 10y agoI tested it on a 680 MB of text. gzip compresses to 246.0 MB. pigz compresses to 245.5 MB. I see similar percent change on a 3.8 MB text file. So they are approximately equivalent.
- toomuchtodo 10y agoWhat was the difference in speed between runs?
- protomok 10y agoYou may be interested in these benchmarks: http://vbtechsupport.com/1614/ http://vbtechsupport.com/1614/
- hrehhf 10y agoWith 2 cores (4 logical) 58.9% decrease in time for the 680 MB file a.txt: # time pigz -c /tmp/a.txt > /dev/null real 0m19.352s user 1m16.148s sys 0m0.344s # time gzip -c /tmp/a.txt > /dev/null real 0m47.093s user 0m46.940s sys 0m0.104s
- kbenson 10y agoFirst, thanks for the numbers, it's useful to see real world examples. Second, and this isn't meant to be a critique (I'm just trying to understand phenomena I see), is there a reason you prefer presenting it as a percentage decrease? Every time I read "X% decrease" I feel obliged to read the source numbers because I'm never sure if the person is using the terminology correctly or not (you are), since so often people mess that up. For myself, I generally use "X ran in Y% of the time Z took." specifically because I don't want people to misinterpret. Is the "X% decrease" presentation preferred/taught, or considered standard? Am I alone in feeling it's more likely to be misinterpreted? (Sorry your comment is the one I brought this up on, I've just been wondering this for a while.)
- teej 10y agoI always state as relative change. I find that people can get really confused if you were to say "X ran in 110% in the time of Y" even though it is stated in a clear way. My preferred way of communicating this concept is "we observed a +10% change in X compared to Y." I always use a +/- sign and this helps signal that I'm talking about a relative change. If I am comparing percents, I'll always specify "relative" or "absolute" change though I prefer to use relative change. Occasionally if the change is small, I will use basis points instead of percents to communicate absolute change.
- pwg 10y agoActually, it does not compress less efficiently, but you do not learn this fact from the README. Instead you have to look inside the man page (https://raw.githubusercontent.com/madler/pigz/master/pigz.1 https://raw.githubusercontent.com/madler/pigz/master/pigz.1): The input blocks, while compressed independently, have the last 32K of the previous block loaded as a preset dictionary to preserve the compression effectiveness of deflating in a single thread.
- spullara 10y agoI'll check it out. However, if you have to wait for the previous block to compress the following block, I don't see how you can parallelize it completely. My assumption is that you would have to shard the file at a higher level and still compress those shards independently. That should get close to the same results as using a sequential compressor but for small file sizes both this effect and Amdahl's law would start to show its head. I suppose you could get around that by not parallelizing compression at some minimum size automatically.
- pwg 10y ago> However, if you have to wait for the previous block to compress the following block, I don't see how you can parallelize it completely. That above is not what the man-page says. The block size for any given session is fixed, so you know the boundaries of each block prior to compression. Each block has a copy of the last 32kbytes of data at the end of the prior block. The algorithm used by gzip compresses by finding repetitive strings in the last seen 32kbyte window of uncompressed data, so there is no compression dependency between blocks, even with a copy of the last 32k of uncompressed data from a prior block being present for the current block.
- spullara 10y agoSo the lossage there is that you don't get to compress and generate that dictionary at the same time. Interesting.
- 10y ago
- bpchaps 10y agoWhy doesn't parallelized gzip get more attention? For larger files, it can be quite a pain to see a multicore machine sit mostly idle. Something like this works pretty well, but I've never seen it done outside of my own stuffs: split -l1000000 --filter='gzip > $FILE.gz' <(zcat large_fille.txt.gz)
- haughter 10y agoHow does the strategy for parallelization of gzip in this project differ from the strategy used for LZMA2 parallelism as implemented by 7z?
- Twirrim 10y agopigz is extremely fast and very capable, plus it's packed up and provided for almost every mainstream linux distribution, and can act as a drop in replacement for gzip as it supports the same flag syntax. On the bzip2 side there is pbzip2, which is also a drop in replacement, http://compression.ca/pbzip2/ http://compression.ca/pbzip2/
- treffer 10y agopbzip2 only does parallel compression. lbzip2 can parallelize compression and decompression. With both pbzip2 and lbzip2 I got errors every now and then while everything worked with bzip2. YMMV. Also note that bzip2 is block based (due to bwt) and thus does not compromise compression ratios like parallelized gzip implantations. At 32 cores you can saturate Gbit links (even with good compression ratios!). I hope we will see some more bzip2 love in the future due to it's perfect fit for parallelism.
- 0x0 10y agoErrors sound deeply concerning. What kind of errors? Is it too risky to use [pl]bzip2 for archival reasons?
- treffer 10y agoWell, errors at a rate of one every few terabyte, and most of the time it was mainly decompression, and always caused program abort iirc. But those have been a few years ago, so my best advice would be: retest. You are testing your archives anyway, right? ;-)
- paxcoder 10y ago>You are testing your archives anyway, right? ;-) Umm no, I handle errors which I expect to be reported to me.
- sambe 10y agoThere is also pbgzip if you are into indexed access (genome data).
- tobias3 10y agoAnd the same for LZMA: https://github.com/vasi/pixz https://github.com/vasi/pixz (it's relatively easy to remember those commands)
- stouset 10y agoDoesn't xz already support threads out of the box, by the --threads flag?
- mnordhoff 10y agoThe first stable version of xz with threading was 5.2.0, released in December 2014. A lot of people are running older packages, or were until recently, and habitually use pixz instead.
- _delirium 10y agoCurrent versions of Debian and Ubuntu ship 5.1.1 fwiw, so it's still quite a large number of people. In fact 5.2.0 doesn't even seem to be in Debian unstable yet. Not 100% sure why, but browsing through the wishlist bug open for years [1], it seems to be due to an unfortunately common reason: even widely used open-source software sometimes has surprisingly few maintainers, in this case seemingly one person, who was maintaining the xz package for years but got busy with other things, and nobody else has picked it up. The good-ish news is that the 5.1.x version now seems to be growing old enough that it's starting to block other packages' upgrades (because their upstream now assumes a newer version), which will probably cause enough other Debian maintainers to notice for the situation to be sorted out. Though I did of course use passive voice in that last sentence, as I wait for "the situation to be sorted out". [1] https://bugs.debian.org/cgi-bin/bugreport.cgi?bug=731634 https://bugs.debian.org/cgi-bin/bugreport.cgi?bug=731634
- 0x0 10y agoThe main purpose of this "pixz" appears to be its chunking of the compressed data so that is it partially decompressible (i.e. random access). "xz" has -T/--threads= already for multithreaded processing (although it does seem like pixz has a different default value "all cores" instead of "1 thread") .
- slashcom 10y agohttp://lbzip2.org/ http://lbzip2.org/ also exists for bzip2. It works really well, especially since bz2 files are split into discreet blocks which can be un/compressed independently.
- lateralux 10y agoI'm using pbzip2 to solve this problem. http://compression.ca/pbzip2/ http://compression.ca/pbzip2/
- Slartie 10y agoI was doing that too, until I tried lbzip2 ( http://lbzip2.org/ http://lbzip2.org/). Does basically the same, also a drop-in replacement, but is even faster by a noticeable amount!
- dmourati 10y agoI learned about pigz in the High Performance MySQL O'Reilly book appendix. I used it, and other techniques there to improve our MySQL backup/restore time by 7x. This in turn won first place at a company hackathon.
- wahnfrieden 10y agoHave you compared with mydumper / myloader?
- LowEntropy 10y agoThis post is a test.
- abhishivsaxena 10y agoAny ideas how this would compare to gzip while on a Microserver? I'm thinking of Atom C2750 bare metal from packet.
- opcenter 10y agoLooks like that has 8 cores? I would imagine it would make a huge difference depending on other loads on the system. I gave it a try on my file server at home that has an AMD E-350 processor with 2 cores and it shaved off a good 42% of the total time: > time gzip -v xubuntu-16.04-desktop-amd64.iso xubuntu-16.04-desktop-amd64.iso: 1.5% -- replaced with xubuntu-16.04-desktop-amd64.iso.gz gzip -v xubuntu-16.04-desktop-amd64.iso 119.07s user 4.66s system 98% cpu 2:05.22 total > time pigz -v xubuntu-16.04-desktop-amd64.iso xubuntu-16.04-desktop-amd64.iso to xubuntu-16.04-desktop-amd64.iso.gz pigz -v xubuntu-16.04-desktop-amd64.iso 128.19s user 6.64s system 184% cpu 1:12.97 total
- LowEntropy 10y agoTest
- dorfsmay 10y agoIf you know that what you compress is alphabetised text, and space is more important to you than time, then please use lzips over gzips. For a parallel version: http://www.nongnu.org/lzip/plzip.html http://www.nongnu.org/lzip/plzip.html
- 0x0 10y agoIs this gunzip-compatible? If not, then as long as gzip remains the only viable/compatible compression mode for HTTP I think advances in gzip compression is still valuable.
- eyelidlessness 10y agoI don't think the comment you replied to suggested that gzip is not valuable, only that there's specialized compression available for a special case (alphabetized data).
- dorfsmay 10y agoNo it's not. Lzip is more useful for storing text long term, or send very large text file as email attachments. Text log file and database dump in text usually generate very large text file, lzip can reduce their size by a factor of 5 to 10, compared to gzips.
- Rapzid 10y agoUsed extensively while a system engineer at a hosting company; 10/10, would use again. Excellent utility if you have the cpu cycles to spare and need to cut time; gzip is almost always your bottleneck. Didn't seem to quite scale linearly, but what does. --rsyncable support too.
- axelfontaine 10y agoPrevious discussion 6 years ago: https://news.ycombinator.com/item?id=1233317 https://news.ycombinator.com/item?id=1233317
- Hydraulix989 10y agoProbably useful for servers more than anything where perf really does matter at scale and content type is gzip.
- teej 10y agoI use this extensively in schlepping data around for analytics work. Any data that moves over the network going into or out of a data store gets compressed with pigz.
- pselbert 10y agoIt's amazing to see what somebody like Mark Adler can achieve when they focus on a specific niche. You often hear about the importance of specialization when positioning a business, but it isn't mentioned as much for individuals. His work serves as an ideal case of specialization—if I wanted to ask an expert for advice on compression, or was looking for an outside expert on compression that is immediately who I'd think of.
- mgerdts 10y agopigz parallelizes compression. I've made changes that make it so that it can parallelize uncompression as well. https://github.com/mgerdts/pigz https://github.com/mgerdts/pigz This feature is present in Solaris 11.3 and later. Actually, it is in some later patches to 11.2 as well. I added it to speed up suspend and resume of kernel zones.
- sshaginyan 10y agoWhy not just use GNU parallel? `parallel gzip ::: file1 file2` I wish there was a standard for this type of stuff. So that an app will check for existing child spawns and act accordingly (IPC).
- maxpert 10y agoI love it! I just wish somebody can make a GPU based compression library (doesn't have to be gzip or bzip), every mobile device today is shipping with a GPU, there are some techniques out there like this (http://on-demand.gputechconf.com/gtc/2014/presentations/S4459-parallel-lossless-compression-using-gpus.pdf http://on-demand.gputechconf.com/gtc/2014/presentations/S445...), but I am still waiting for a solid implementation.
- yzh 10y agoYou may find work from my colleagues interesting: Parallel Lossless Data Compression on the GPU: http://www.idav.ucdavis.edu/publications/print_pub?pub_id=1087 http://www.idav.ucdavis.edu/publications/print_pub?pub_id=10... Fast Parallel Suffix Array on the GPU: http://escholarship.org/uc/item/83r7w305 http://escholarship.org/uc/item/83r7w305 I think built on top of the second paper (suffix array paper), there should be fast compression library. We have one implementation in our library CUDPP 2.0, you can try it out if you like: https://github.com/cudpp/cudpp/blob/master/src/cudpp/app/sa_app.cu https://github.com/cudpp/cudpp/blob/master/src/cudpp/app/sa_...
- ganeshkrishnan 10y agoPied Piper already made it.
- snark42 10y agoDoesn't transferring all the data to/from the GPU kill any potential speedup the GPU may offer?
- discreditable 10y agoYou can also use 7-zip for multithreaded gz/xz/bz2/7z/zip compression.
- donatj 10y agoHuh. I was living under the incorrect assumption that gzip was inherently single threaded
- lmeyerov 10y agoWe use this in production at Graphistry, super useful for latency-sensitive dynamic media applications! (In this case, we needed to add an auto-tuner + node bindings.)
- noipv4 10y agoFor bzip2 lovers the tool is pbzip2 ;)
- axelfontaine 10y agoThe big challenge seems to be parallel gunzip as that requires a special stream with an index to work, with no general purpose solution available so far.