3 ms·
As ubiquitous as zlib is, I wonder why it’s not implemented in hardware like many audio and video codecs are.
by dharbin 5y ago
As ubiquitous as zlib is, I wonder why it’s not implemented in hardware like many audio and video codecs are.
- wbl 5y agoIt can be. But searching a large dictionary for the longest match is a pain. Audio and video codecs tend to have a much smaller working set or lots of parallelism for hardware depending on what stage of the process they are in.
- erk__ 5y agoIt is implemented in hardware on IBM's mainframes, although not upstreamed in the main zlib repo [0]. It gives a massive speedup in my experience up to 20x for decompression and 220x (!) for compression [0]: https://github.com/madler/zlib/pull/410 https://github.com/madler/zlib/pull/410
- wolf550e 5y agodeflate is implemented in hardware on game consoles. PS5, in addition to zlib, has hardware [0][2] decompressor for RAD Oodle Kraken [1], which is of course much better than zlib. If at all possible, especially if you control both ends, please use zstandard[3] instead of zlib/deflate/gzip/zip for free general purpose compression. It is more advanced by about 30 years of technology development. It always spends less time to produce smaller files which are then decompressed insanely fast. There is no tradeoff, it is always better. It has a very wide range of compression levels you can choose from if the good default doesn't fit your need. It also supports multithreading, custom dictionaries and long range compression out of the box. 0 - https://www.tweaktown.com/news/71340/understanding-the-ps5s-ssd-deep-dive-into-next-gen-storage-tech/index.html https://www.tweaktown.com/news/71340/understanding-the-ps5s-... 1 - http://www.radgametools.com/oodlekraken.htm http://www.radgametools.com/oodlekraken.htm 2 - https://twitter.com/rygorous/status/1240341184867758085 https://twitter.com/rygorous/status/1240341184867758085 3 - https://facebook.github.io/zstd/ https://facebook.github.io/zstd/
- mananaysiempre 5y ago> There is no tradeoff, [Zstd] is always better. Well of course there’s the tradeoff in memory usage, utterly irrelevant on a desktop and probably even on a server, but you’re never decompressing a Zstandard file (compressed with standard options) on an STM32 (36 to 144 MHz ARM Cortex-M) micro or a similar RAM-impoverished but not entirely wimpy system. That is no secret, though, IIRC part of the original motivation for Zstandard was that modern hardware makes old window size choices obsolete. A less obvious point is that a modern implementation of Deflate is, if not necessarily always faster, not as catastrophically slow compared to Zstandard as Zlib[1]. (In other benchmarks I’ve seen libdeflate be something like half again as slow as Zstandard, but at least it’s not multiple times slower.) [1] https://lemire.me/blog/2021/06/30/compressing-json-gzip-vs-zstd/ https://lemire.me/blog/2021/06/30/compressing-json-gzip-vs-z...
- user-the-name 5y agoIf you are working on a microcontroller, odds are you will be keeping your entire file in RAM as that is all you have, in which case you should not need any memory for a window at all, as you can just use the file itself. I don't know if the zstd implementation supports working like this, though.
- wolf550e 5y agoI believe you can compress into the source buffer and keep the read offset ahead of the write offset, but I don't think you can decompress this way.
- user-the-name 5y agoThis is about the window buffer, not the source and destination buffers. They are still distinct, but you can combine the uncompressed buffer and the window buffer, if you have, or are going to have, the entire uncompressed data in memory.