3 ms·
FWIW, llvm-mca estimates 448 clock cycles per 100 iterations of the AVX2 loop vs 528 cycles for the AVX512 loop with `-mcpu=cascadelake`. That suggests the AVX5
by celrod 5y ago
FWIW, llvm-mca estimates 448 clock cycles per 100 iterations of the AVX2 loop vs 528 cycles for the AVX512 loop with `-mcpu=cascadelake`. That suggests the AVX512 loop should be about 2*(448/528)=1.85 times faster.
- pbsd 5y agollvm-mca is highly unreliable when it comes to AVX-512. It thinks 3 512-bit vpaddd, vpsubd can be run per cycle. Adjusting for that you get 622 cycles instead of 528.