8 ms·
The Intel algorithm should be much faster. The slicing algorithms are limited by load unit bandwidth, which appears to be 2 loads/cycle on Zen 2, while PCLMULQD
by ack_complete 3y ago
The Intel algorithm should be much faster. The slicing algorithms are limited by load unit bandwidth, which appears to be 2 loads/cycle on Zen 2, while PCLMULQDQ should be able to process 16 bytes every 2 cycles (8 bytes/cycle peak). 4.5GB/s is about 1-1.5 bytes/cycle. zlib-ng has an implementation.
- mxmlnkn 3y agoThat sounds more than I, for some reason, expected. I'm not sure why, but I was expecting "only" 2x speedup at maximum, while the CRC32 adds "only" 5% overhead. But, 5-8x speedup would be really nice. With that, the the CRC32 overhead would become immeasurable. I'll definitely have to look into it now.