4 ms·
Without digging into it too much, I would bet that most of the improvement from ICC comes from their awesome implementations of transcendental functions used in
by dsharlet 10y ago
Without digging into it too much, I would bet that most of the improvement from ICC comes from their awesome implementations of transcendental functions used in this program (exp, sqrt, etc.).
Still, the general point is good. last time I tried ICC on a heavily vectorized (with intrinsics) program, ICC was a 30% boost over clang or gcc.
- nkurz 10y agoThis was my first guess, but as I dig a bit deeper I'm less sure. It looks like the difference is the implementation of log and exp, but both of them are linked to the same libm from glibc. The difference is that Intel is inlining an AVX2 optimized version that uses FMA, but Clang and GCC are using the compiled function from libm. The Ubuntu packages I'm using are compiled appear to have been compiled to support AVX, but not AVX2, and hence no FMA. So while GCC and Clang are much slower, this is really an issue of having suboptimal system libraries. The speedup is real, but it's not clear that either GCC or Clang are actually to blame.
- slavik81 10y agoNot an expert on this, but I think the compiler has to inline the function for that to work well. No matter how good the system library implementation is, the function signature is defined one scalar float, which limits vectorization.
- pcordes 10y agoThat's correct; I'm sure icc has vectorized versions of those functions to inline. Each monte-carlo iteration is independent, so the code vectorizes very easily if you have a vectorized RNG and vectorized exp() and log(). (vector sqrt() is available in hardware, so even gcc and clang will vectorize that, but not exp / log even with -ffast-math)
- nkurz 10y agoOK, I was able to test this more rigorously. Yes, the vast majority of the difference between Intel and GCC is just the better implementations of exp() and log(). These are not vector implementations that calculate multiple results simulataneously, just scalar implementations that make better use of the available instruction set. When I can get g++ to use Intel's libimf instead of glibc's libm, it's only ~5% slower than icpc. I'd guess much of the remaining difference is function call overhead vs inlining. Clang still lags by another 10%, but I think that's probably because I haven't figured out quite the right incantation to get it to use libimf without also disabling some other useful optimization. Here are the commands that I ended up with: clang++ -fno-finite-math-only -march=native -Wall -Wextra -g -O3 option.cc -o option -Wl,-rpath=/opt/intel/compilers_and_libraries/linux/lib/intel64 -L/opt/intel/compilers_and_libraries/linux/lib/intel64 -limf -lintlc g++ -fno-finite-math-only -march=native -Wall -Wextra -g -Ofast option.cc -o option -Wl,-rpath=/opt/intel/compilers_and_libraries/linux/lib/intel64 -L/opt/intel/compilers_and_libraries/linux/lib/intel64 -limf -lintlc icpc -march=native -Wall -Wextra -g -Ofast option.cc -o option The "-fno-finite-math-only" disables an otherwise good clang++ and g++ optimization that requires a __finite_exp() function that libimf does not have. clang++ seems to also need to switch from -Ofast to -O3 to make this stick. What I don't know yet is whether recompiling glibc (or upgrading to the most recent glibc) will produce better performance out-of-the-box on recent Intel. My searches aren't turning up much information --- anyone know? gcc and clang will vectorize that, but not exp / log I don't know how well it's currently working, but it looks like this might have changed recently: https://sourceware.org/glibc/wiki/libmvec https://sourceware.org/glibc/wiki/libmvec