4 ms·
I've been working on LoopVectorization in Julia, and benchmarking against a few compilers. Intel's compilers are far ahead of GCC and LLVM in vectorizing loops.
by celrod 7y ago
I've been working on LoopVectorization in Julia, and benchmarking against a few compilers.
Intel's compilers are far ahead of GCC and LLVM in vectorizing loops.
LLVM (even with Polly) fairs worst in my benchmarks, so it does not look like a future where Intel is obsolete is on the horizon yet.
However, I am excited for Flang and its FIR MLIR-dialect. I haven't benchmarked MLIR-optimized code at all. I'm sure that will change things, but until I test I have no idea by how much.
- noobermin 7y agoI'd say it's more than merely far ahead. When it's your hardware and you've been doing HPC optimization for decades, you are leagues ahead.
- stephencanon 7y agoThere are no great secrets to the HW or vectorization. It’s mostly a question of willingness to devote resources to the problem and hire folks (e.g. Aart Bik, who literally wrote the book on vectorization at Intel and is now working on the TF compiler team at Google and contributing to LLVM).
- gnufx 7y agoThis "leagues ahead" is simply not true experimentally. I ran the Polyhedron benchmarks on SKX with profile-guided optimization. Pre-release gfortran 10 was (insignificantly) faster on the bottom line than beta ifort (from oneapi). That was reversed for gfortran 8 v. ifort 18. It's also not true that HPC performance is generally dominated by code generation rather than libraries and communication costs, but obviously mileage varies.
- stephencanon 7y agoNote that a lot of Intel’s advantage in SIMD, historically, came from optimization modes that simply (unsafely) assume that there are no dependencies that would block vectorization (instead of doing the conservative analysis like GCC and Clang). It makes for impressive benchmarks, but is borderline-unusable in much real code. Intel’s compiler team has actually suggested some patches adding the same mode to LLVM, though I’m not sure what the current status is, since the initial reaction was not overwhelmingly positive.
- celrod 7y agoInteresting. Would this be safer in a language like Fortran, where (without aliasing between separate arrays), loop dependencies should be more obvious? I think it'd be nice to be able to activate this mode through pragmas. Does "#pragma omp simd" result in more aggressive use of blocking?
- gnufx 7y agoI don't know how omp simd is related to "blocking", but it can slow code by a factor of two compared with GCC optimizing the loop normally, because you get AVX(51)2 but not FMA. (Observed with a generic C GEMM, which got about 60% of the micro-optimized one after removing such pragmas and just letting gcc do its -Ofast thing.)
- mkj 7y agoMy experience in optimising c++ and Fortran with Intel compiler 16 is that it is safe regarding dependencies. Often #pragma ivdep is required to vectorise loops it can't accept otherwise.
- gnufx 7y agoThere will be cases for which it's not so, but every time recently that I can remember someone saying ifort/icc vectorizes some numerical code profitably, and gfortran doesn't, and I had the code, I got GCC vectorizing it with equivalent flags. ifort defaults to incorrect behaviour for floating point (something like gcc -ffast-math). [Edit: As I now see it says in a sibling.]