3 ms·
Recent versions of GCC as well as the Intel compilers can autovectorize with the right flags (-ffast-math with GCC, -fast with Intel), and yield better performa
by celrod 7y ago
Recent versions of GCC as well as the Intel compilers can autovectorize with the right flags (-ffast-math with GCC, -fast with Intel), and yield better performance with the <math.h> functions than you'll get from SLEEF on x86_64.
LLVM doesn't have a vector math library, so SLEEF could help you there.
If you're using single precision and AVX512, a >10x speed increase is likely. Otherwise, you'll probably get less than that.
These functions are very accurate, most to withing 1 ULP of least precision. That is, if the answer they provide isn't the correctly rounded floating point answer, then it'll be either the next or previous representable floating point number.
There is a lot of room for giving up accuracy in the name of speed (eg, using less terms in the polynomials).
- svantana 7y agoDo you have a benchmark showing that gcc is faster than sleef? Sorry I wasn't clear, by "10x" I meant 10x faster than the standard library with the fastest compiler options (-O3 -ffast-math -avx512), but I've only tested with clang.