3 ms·
Looks interesting, but my main objections to general adoption the same as bzip2, lzma and context modelling based codecs - decompression speed. Compressing log
by klauspost 4y ago
Looks interesting, but my main objections to general adoption the same as bzip2, lzma and context modelling based codecs - decompression speed.
Compressing logs for instance, decompression speed of 23MB/s per core, is simply too slow when you need to grep through gigabytes of data. Same for data analysis, you don't want your input speed to be this limited when analysing gigabytes of data.
I am not sure how I feel about you "stealing" the bzip name. While the author of bzip2 doesn't seem to plan to release a follow-up, I feel it is bad manner to take over a name like this.
- vintermann 4y agoWith the other parts of the codec, I doubt it's possible for files compressed with this, but one of the strengths of BWT-based compression is that there has been a lot of research on search operations directly on compressed data.
- perihelions 4y agoThat's funny, 23 MiB/s is exactly what I get for reading systemd logs (on an NVME SSD). Is it supposed to be otherwise? $ sudo journalctl -r | pv -a > /dev/null [22.8MiB/s]
- yakubin 4y agoThat appears to be systemd being slow. $ dd if=/dev/urandom of=test bs=1G count=1 iflag=fullblock $ gzip -k test $ zcat test.gz | pv -a >/dev/null [ 228MiB/s] $ sudo journalctl -r | pv -a >/dev/null [13.1MiB/s] UPDATE: Gzip with more real-world data[1]: $ gzip -k adventures-of-huckleberry-finn.txt $ zcat adventures-of-huckleberry-finn.txt.gz | pv -a >/dev/null [ 151MiB/s] [1]: <https://gutenberg.org/files/76/76-0.txt https://gutenberg.org/files/76/76-0.txt>
- klauspost 4y agoBy "compressing" random data you are bypassing gzip, since it will just store your data as uncompressed blocks, making "decompression" a memcopy. With real data, deflate maxes out somewhere around there either way, but that is a bit coincidental. With modern CPUs getting increasingly smaller IPC improvements this will likely be pretty much the max decompression speed we can expect from gzip going forward.
- yakubin 4y agoIt actually made the file bigger (1.1G from 1.0G) :) I was getting the same numbers with text data I have scattered on my disk, but those were small, so I decided to generate a bigger file. But, yes, I agree a more robust benchmark would use a Mark Twain novel e.g.
- erk__ 4y agoI think the lack of speed here is more that it has to serialize the data from disk into a readable format. I assume using the `--grep=` option is faster than piping it through grep because of this
- perihelions 4y ago--grep= That's exactly what I needed to know! I'm glad I asked the stupid question. Thank you!
- bayindirh 4y ago> I am not sure how I feel about you "stealing" the bzip name. While the author of bzip2 doesn't seem to plan to release a follow-up, I feel it is bad manner to take over a name like this. I think it boils down to the feelings of the author (of the previous format). I don't think PKWARE feels bad because ZSTD is a homage to ZIP. Similarly if someone created a follow-up file format to something I've designed, I'd just want credit or a link to my version as a homage and pointer for history continuity, nothing else. Open source software is designed to be mangled, modified, shared and leapfrogged. If a completely different implementation advertises itself as a newer iteration of a format because it's built on the same theory, I think it's ethical if the developer is not intending to capitalize name. Either case, if the original developer returns to the game, it can create a BZIP4 and points to the diversion as, "hey, somebody liked BZIP2 too much and created this, give him/her a kudos", and continue.
- altairprime 4y agoPkware might not have been so forgiving if someone had released ZIP2. Incrementing the version number like that is only an acceptable thing in relatively unusual circumstances, but does happen sometimes; and still, I would really hesitate to say that it’s a good idea for a third party to call itself bzip3. The original author replaced their own bzip release with bzip2 to avoid a patent issue with arithmetic coding, so this is the first time a third party has done so: https://web.archive.org/web/19980704181204/http://www.muraroa.demon.co.uk/ https://web.archive.org/web/19980704181204/http://www.muraro... So if the release of bzip3 is approved by the current maintainer, then I guess it’s fine, but otherwise it makes me uncomfortable to consider using under this name.
- eru 4y ago> Open source software is designed to be mangled, modified, shared and leapfrogged. I agree in spirit, but I can also see why someone might want their source to be free and mangleable, but still care about trademarks. (Just imagine Linus Torvalds getting lots of emails with support requests for a hypothetical Linux2 operating system that I wrote, and that he has no relation with. That could become pretty annoying; even if he doesn't mind me taking his source code.)
- usefulcat 4y agoIf you’re using xz, pixz can do multithreaded decompression. It’s still xz/lzma, so still expensive to decompress, but at least that allows you to throw as many cores as you want at it.