4 ms·
Nice work! How much manual SIMD are you using?
by ComputerGuru 3y ago
Nice work! How much manual SIMD are you using?
- mxmlnkn 3y agoCurrently, none. ISA-L might use some, I'm not sure. I have also tried to accelerate some crucial algorithms with SIMD and my trial and errors are still inside the src/benchmarks subfolder. But, in all of those cases, a simple std::string_view::find, std::transform, or lookup tables turned out to be equally fast or faster and, of course, is more portable. Even the CRC32 algorithm uses "simple" lookup tables (slice-by-16) and avoids any SIMD while still achieving ~4.5 GB/s per core. However, there is some Intel Whitepaper [0] showing how to use PCLMULQDQ for CRC32, which might not be much faster but it would reduce cache pressure. AVX-512 even has a VPCLMULQDQ. [0] https://www.intel.com/content/dam/develop/external/us/en/documents/clmul-wp-rev-2-02-2014-04-20.pdf https://www.intel.com/content/dam/develop/external/us/en/doc...
- ack_complete 3y agoThe Intel algorithm should be much faster. The slicing algorithms are limited by load unit bandwidth, which appears to be 2 loads/cycle on Zen 2, while PCLMULQDQ should be able to process 16 bytes every 2 cycles (8 bytes/cycle peak). 4.5GB/s is about 1-1.5 bytes/cycle. zlib-ng has an implementation.
- mxmlnkn 3y agoThat sounds more than I, for some reason, expected. I'm not sure why, but I was expecting "only" 2x speedup at maximum, while the CRC32 adds "only" 5% overhead. But, 5-8x speedup would be really nice. With that, the the CRC32 overhead would become immeasurable. I'll definitely have to look into it now.
- doophus 3y agoFor portable SIMD, have a look at ISPC. It allows you to write a function once, compile it for multiple instruction sets, and then automatically select the best one to use at runtime. You don't get the precision of hand-crafted SIMD, but it can grant some easy wins!