7 ms·
Intel QuickAssist Technology Zstandard Plugin for Zstandard
- jeffbee 3y agoI love the QAT libraries and I feel their abilities are overlooked. Intel also has the igzip library that does not even require QAT and it radically faster than zlib, which is handy in older applications where gzip is unavoidable despite its obsolescence. The major downside of course is it is quite tricky to use this stuff in practice. In the cloud, you need a bare metal instance that exposes the QAT peripheral, and they are relatively scarce. And this whole generation of hardware is only just beginning to land in public clouds. For machines you own, you will need to scrutinize Intel's somewhat ridiculous product matrix in order to acquire a Xeon that has QAT.
- VWWHFSfQ 3y ago> older applications where gzip is unavoidable despite its obsolescence Is gzip actually obsolete or are there just newer alternatives? gzip is still everywhere
- jeffbee 3y agoThe zlib implementation is not optimal in any way that I am aware of. You can get better compression with the same compute time, and you can get the same compression with less time, in every application I have seen, by using some other thing. I usually want to look at zstd, lz4, and brotli in addition to zlib. That said, the igzip re-implementation is really good. If you can use it, igzip moves gzip closer to the optimal frontier.
- rincebrain 3y agoThere's also the various zlib forks by cloudflare et al, if you want to go down that road. But yes, lz4 and zstd would be what I'd recommend people use.
- dralley 3y agoThe code quality of some of the cloudflare changes is questionable. Optimizations aren't cleanly separated by commits, some of the licensing is dicey, etc. zlib-ng tried to merge as much as they could from the cloudflare fork without getting into licensing issues.
- rincebrain 3y agoOh, sure, it's all exciting in that respect from some of those, but I was more pointing to the collection of forks as a whole as a thing to evaluate if you have to look at zlib. I would just go for lz4 and zstd, really, at this point, but if for some reason gzip is necessary or fits your use case better, ... (Though I'd really like to hear what that is.)
- jandrese 3y agoThe only place gzip wins is in compatibility. Everything supports gzip. zstd, lz4, brotli, and others aren't included by default in most places yet, although the situation is starting to improve. Plus, if you want super compatibility there is older-than-dirt pkzip/infozip.
- benatkin 3y agoit's optimal in that it's a community project and not an extremely unethical BigCorp project
- wolf550e 3y agoIt's obsolete. DEFLATE format is limited to 32KB LZ window with huffman coding. Zstd can use a much larger window (8MB recommended) and a much better entropy coder: https://github.com/Cyan4973/FiniteStateEntropy https://github.com/Cyan4973/FiniteStateEntropy
- klauspost 3y agoLast I checked QAT was limited to 64KB backreferences and independent 128KB blocks. Now you know why they are comparing 16KB payloads only.
- terrelln 3y agoYeah, that is definitely a limitation with QAT. It isn't a great fit for larger data, as it can quickly lose compression ratio due to its smaller window. However, there is a lot of compression done on data that is <64KB. E.g. compression in RocksDB. We also have some half baked ideas to combine a fast SW match finder that only looks for matches >64KB away, and supplements the matches that QAT finds.
- andrius4669 3y agoany idea how igzip's non-QAT path compares to zlib-ng[1]? [1] https://github.com/zlib-ng/zlib-ng https://github.com/zlib-ng/zlib-ng
- jeffbee 3y agoIt's fairly easy to check for yourself, but with local builds on a desktop CPU without QAT, compressing a 150MB JSON input, with igzip and minigzip (from zlib-ng) both producing a smaller output than gzip, gzip needs 1.187s, minigzip needs 0.665s, and igzip needs 0.313s.
- powturbo 3y agoYou can use TurboBench [1] to benchmark igzip, zlib-ng and others. Download TurbBench from Releases [2] Here Some Benchmarks: - https://github.com/zlib-ng/zlib-ng/issues/1486 https://github.com/zlib-ng/zlib-ng/issues/1486 - https://github.com/powturbo/TurboBench/issues/43 https://github.com/powturbo/TurboBench/issues/43 [1] https://github.com/powturbo/TurboBench https://github.com/powturbo/TurboBench [2] https://github.com/powturbo/TurboBench/releases https://github.com/powturbo/TurboBench/releases
- pizza 3y agoThis is great, I can think of several places where zlib is the bottleneck and would love to speed them up
- selectodude 3y agoIntel’s software is simply best in class (n=2).
- brirec 3y agoQAT does support SR-IOV on platforms with IOMMU support, so you can at least delegate access to QAT across multiple VMs. I don’t think any public cloud providers offer this, though.
- berbec 3y agoA massive power and CPU load decrease in Zstandard is a big win for Intel. AMD has been racking up major pluses in the enterprise space, with RAM, PCIE and core count advantage. Showing that any Intel is faster at such a major CPU load task is a big deal. That's not to detract from everything AMD has done, but hardware is only the first step. Software that properly uses the features your hardware provides is just, if not more, important. I love the fact AMD is pushing Intel so much. Pre-C2D days were amazing because we had two vibrant, innovative companies pushing to the edge of possible; trying to out-do each other. Pre-Ryzen was a horrible time. Do you want to spend $500 to upgrade from a 4-core intel 4000-cpu to an intel 5000-cpu? You'll get DDR4 and 1% IPC. Now we get massive IPC, clock speed, ram and PCIE improvements on a regular basis. Competition is great, especially for the consumer.
- jiggawatts 3y agoFrom the benchmarks I’ve seen, QAT is a desperate attempt to gain some competitive advantage over AMD. It’s not clear if there is any actual advantage for real workloads. Sure, benchmarks of pure compression will come out ahead for Intel, but a typical sever running something like a database engine with compressed disk storage is likely a win for AMD. The new EPYC 9004 chips are much cheaper than the Intel CPUs with QAT, and they have more cores, and much higher clock speeds. It’s hard to believe there’s any scenario where the Xeon CPUs provide better overall value…
- nvm0n2 3y agoHow is compression not a real workload? Zstd is an excellent codec that's rapidly proliferating - for example Chrome is integrating it at the moment. https://bugs.chromium.org/p/chromium/issues/detail?id=1246971 https://bugs.chromium.org/p/chromium/issues/detail?id=124697... Relatively few servers run databases. Many more are running web servers and given the huge efforts web devs go to in order to optimize response sizes, much faster, lower power and lower latency compression seems like an instant win. All that's required is for web servers to integrate the zstd library and this QAT thing, and that can be enough to tip the balance especially for IO bound servers that are mostly just doing string interpolation and waiting for backends. And you mention a DB with compressed disk storage. How is better compression not a huge win for that use case? Databases are usually disk IOP, latency and CPU power constrained, and this is a win for all of those cases. Finally, clearly this tech can be applied to other algorithms not just zstd. Presumably they highlight zstd because the library is actively maintained and was willing to add this (probably single use) "plugin" API. Really they should have just integrated it directly instead of complicating things for every zstd user, but I guess it adds extra dependencies or increases code size or something. The wins are big enough that other codec libraries will probably use the same approach and eventually it'll be abstracted by APIs that are less code size sensitive. Yes, AMD is doing great right now, but Intel have had an edge when it comes to specialized CPU features for a long time. For example AMD are only just now catching up to where Intel were with SGX 8 years ago.
- scrubs 3y agoChecking this out tomorrow. Side work project is better compressing 10s of TBs of files.
- soulbadguy 3y agoPlease report back your findings, I am sure I am not the only one wanting to see some real world results
- throwaway81523 3y agoI'd like to know what QAT actually is? Some kind of special purpose hardware extension for LZ and crypto? I used to compress a lot of files using xz and it is slow, but not THAT slow. I have a now-ancient 4 core i7 server and it takes a day or so hours to compress a TB of text. I haven't looked at it in a while and don't remember actual numbers, but the point is that if you don't need realtime response, the cpu usage of these algorithms is tolerable without special hardware.
- jeffbee 3y agoIt's a PCIe device that is either on a stick (earliest generation), integrated into the platform chipset, or integrated into the CPU (latest generation). Nobody outside Intel really knows what's in it, but you submit requests to it and it encrypts, decrypts, checksums, or compresses something. If this sounds familiar it is because we had stuff like that in the 90s.
- jevinskie 3y agoThere is some enum or #define in a header I found recently that suggested there was once QAT acceleration of something XML related! I can’t find any documentation about it though. edit: CPA_ACC_SVC_TYPE_XML https://github.com/ravynsoft/ravynos/blob/ee81203faa28ab3e6a987b1e2c4f3b759d02ebfa/sys/dev/qat/qat_api/include/cpa.h#L419 https://github.com/ravynsoft/ravynos/blob/ee81203faa28ab3e6a...
- bitbckt 3y agoFinally. I’m tired of seeing zlib in QAT reviews alone - it’s largely irrelevant to situations where I might want to choose Intel (for QAT) over AMD. I don’t fault Intel for choosing web frontend acceleration over storage first, but this has been a long time coming.
- pclmulqdq 3y agoQAT is awesome, and flips the script on the notion of core count primacy. In many servers, QAT when properly used will save several CPU cores, since cores spend a lot of time compressing and encrypting stuff. However, the software layer has always been Intel's weakness, and I'm not entirely sure they got this one right.
- bluedevilzn 3y agoIs this available in any public cloud?
- wmf 3y agoAWS has it.
- c_o_n_v_e_x 3y agoOut of curiosity, what are some use cases? The only one I'm even remotely familiar with is https encryption offloading.
- wmf 3y agoThis sounds like it would be great for filesystem compression (e.g. ZFS).
- nvm0n2 3y agoIntel's software has always been pretty good for me. You're complaining about the need for a driver, right? That seems reasonable, more like a weakness of cloud computing that it takes so long for clouds to expose the real capabilities of the hardware. It's a win for those who can (re)master dedicated machines and attendant Linux sysadmin issues.
- tecleandor 3y agoI think the driver and software resources have been very messy. There are different versions, hardware revisions, different QAT generations... and at first glance is very confusing to know how you can run certain accelerations. For example, the Zstandard plugin has this requirements: Hardware Requirements Intel® 4xxx (Intel® QuickAssist Technology Gen 4) Software Requirements ZSTD* library of version 1.5.4+ Intel® QAT Driver for Linux* Hardware v2.0 What's Gen 4? (I later found that means '4th Generation Intel® Xeon® Scalable Processors'. But not any Gen4, you have to find one with QAT included, and the product matrix is huge.) What's the v2.0 of the hardware? (I'm not completely sure about that) Does it automatically accelerate zlib and openssllib calls, or do you have to patch them? (I think you have to patch them) I think they're stabilizing the API from now on, but it's not as simple as buying a CPU or one of the older PCIe cards, and loading the driver.
- estebarb 3y agoIs there an easy way to use these accelerators from the CLI? Sometimes I have to decompress several TBs of gzip files, but I don't want to rollout my own decompressor in C. I know that Graviton 2 includes a compression accelerator as well, but no idea how to use it (easily).
- xxs 3y agoThe accelerator (at least this particular use) won't help decompression.
- mxmlnkn 3y agoIf you have large gzip files to decompress, I can recommend igzip[1] for fastest single-core decompression or my pet project rapidgzip[2] for effective multi-core decompression, both have simple CLI tools. For using accelerators (QAT, I assume, because the article is about this) from the CLI, then QATzip[3] comes with a command line tool qatzip which can be used as described in the project ReadMe. I didn't test it, though, as I have no QAT-enabled device. [1] https://github.com/intel/isa-l/tree/master/igzip https://github.com/intel/isa-l/tree/master/igzip [2] https://github.com/mxmlnkn/rapidgzip https://github.com/mxmlnkn/rapidgzip [3] https://github.com/intel/QATzip#test-qatzip https://github.com/intel/QATzip#test-qatzip
- metta2uall 3y agoInteresting that Intel's code for this includes numerous references to LZ4, as if that's the actual algorithm the hardware originally aimed to accelerate.. So seems like LZ4 and ZSTD are quite similar? https://github.com/intel/QAT-ZSTD-Plugin/blob/main/src/qatseqprod.c https://github.com/intel/QAT-ZSTD-Plugin/blob/main/src/qatse...
- lifthrasiir 3y agoIt's actually LZ4S, which is an intermediate format very similar to LZ4. LZ4 can be regarded as an LZSS-only compression format with no further coding, so any format that uses LZSS as its core modelling scheme will benefit from LZ4S acceleration. (QAT also has specific support for DEFLATE and Huffman coding, but that's all.)
- xxs 3y agoPretty much all LZ77 class of algorithms try finding longest matches in a dictionary. (both zstd, lz4... and deflate [(g)zip] are all lz77)
- sanqui 3y agoBtrfs, the file system I use, utilizes zstd for transparent compression. That's using a lot of CPU all the time on laptop. So more efficient compressing is great news! Is this for future CPUs?
- pgtan 3y agoWell, POWER9 and later CPUs have builtin hardware compression using the 8-4-2 algorithm. https://en.wikipedia.org/wiki/842_(compression_algorithm) https://en.wikipedia.org/wiki/842_(compression_algorithm) https://github.com/plauth/lib842 https://github.com/plauth/lib842
- loeg 3y agoGood to see more in the compression offload space. Several years ago we ended up running a custom gzip softcore on an FPGA (I believe) co-located on a NIC to get somewhat better gzip compression performance than software. (We were pretty short on PCIe physical capacity in that model.) Dealing with the gzip core vendor and the FPGA vendor (both in wildly different timezones) was a little unpleasant.
- SilverBirch 3y agoThis may be a really dumb question... but: Is this transparent? Like, can I compress some data using QAT to create a zstd file, email it to my friend and have them decompress it without QAT? From the way that this is described it sounds like they're replacing the sequence producer, but presumably that doesn't matter as long as the format you encode those sequences adhere to some standard format?
- ahofmann 3y agoWhy do they show different compression levels in their graphs? That seems kind of fishy to me.
- pxeger1 3y agoLooks like they choose the levels to make the compression ratios roughly equal. > For the Silesia corpus, data compression ratios are: QAT-ZSTD level 9: 2.76 zstd level 4: 2.74 zstd level 5: 2.77 This presumably also means the best possible compression that QAT can achieve is worse than what vanilla zstd can do.
- klauspost 3y ago.. for small payloads. QAT cannot compress across blocks. So as soon as your input is more than 128KB the compression ratio tanks. Usually well below what even level 1 of software zstd does. They choose the input very carefully.
- mastax 3y agoInteresting. That sounds basically perfect for ZFS where the default is 128KB blocks and zstd compression.
- baybal2 3y agohttps://ark.intel.com/content/www/us/en/ark/products/125200/intel-quickassist-adapter-8970.html https://ark.intel.com/content/www/us/en/ark/products/125200/... At "only" 100GBPs per adapter, with said adapter costing like 1 Epyc, does it make sense? Epycs can do 400gbps of compression in software, without much SSE, and handwritten assembler.
- dale_glass 3y agoI think that's a very welcome improvement. With NVMes that go at 7 GB/s we're now at the point that it can be hard to do anything useful with the data fast enough. So I think good acceleration for things like compression is going to be a big help.
- NelsonMinar 3y agoQuickAssist Technology is new to me. What hardware supports it? A quick look suggests it's just a few Xeon processors or else a $800 peripheral card. Is it at all related to "Quick Sync", the name for the video compression acceleration in newer Intel CPUs?
- wmf 3y agoQAT and Quick Sync are unrelated.
- stefantalpalaru 3y ago[dead]
- truth_seeker 3y agohah ! Nailed it. Best way to optimise the reusable software is to turn it into single hardware CPU instruction