3 ms·
A parallel PCLMULQDQ version is faster than a serial CRC32. Because CRC32 has a latency of 3 cycles. A parallel CRC32 can be slightly faster with more data seg
by illumen 11y ago
A parallel PCLMULQDQ version is faster than a serial CRC32. Because CRC32 has a latency of 3 cycles.
A parallel CRC32 can be slightly faster with more data segment used, and more code size.
Also PCLMULQDQ can use an polynomial I think. So they could make a compatible version.
The lookup table version they used has to have used lots more CPU cache. Also, I would suggest they look at reducing the amount of times they call the function. Since there are two ways to speed up functions that are called a lot. One is to stop calling it so many times.
Or since IO is limiting, they could consider compressing/decompressing the data as well. Using something fast like LZ4 compression. This will give them faster IO and a checksum at the same time. Something like Blosc can be faster than memcpy. Especially since the data is replicated over the network as well, and they would save space on their storage.
Modern CPU performance optimization is very often about memory IO. If the data is in L2/L3 then you can do a LOT of computation on it, almost for free, compared to the time it takes to get it into L2/L3.
During that time waiting for memory/disk/network IO, they might consider other things. Like indexing, better checksums (like SHA1) or even encryption. Especially if the data is already in L2/L3.