3 ms·
The floating point comment leaves out that one can use #pragma omp simd reduction(+:res) as a more precise way to achieve vectorization in the reduction (
by jedbrown 6y ago
The floating point comment leaves out that one can use
#pragma omp simd reduction(+:res)
as a more precise way to achieve vectorization in the reduction (compile with -fopenmp-simd to only use it for SIMD without linking an OpenMP library): https://godbolt.org/z/17oTz1 https://godbolt.org/z/17oTz1
Unfortunately, the pragma is not supported with the new-style class iterators in a released compiler, though it works in clang-trunk: https://godbolt.org/z/hbP11W https://godbolt.org/z/hbP11W
Note that Clang disables floating point contraction by default (so no vfmadd instructions), despite them being more accurate. One usually wants this globally (-ffp-contract=fast) except when trying to bitwise reproduce software compiled for pre-Haswell.