4 ms·
There are unfortunately a lot of serial guarantees that are unavoidable even if compilers could optimize perfectly. For example: 'for(double v: vec) sum+=v' Fl
by kolbe 3y ago
There are unfortunately a lot of serial guarantees that are unavoidable even if compilers could optimize perfectly. For example: 'for(double v: vec) sum+=v'
Floating point addition is not associative, so summing each value in order is not the same as summing every 8th element, then summing the remainder, which is how SIMD handles it. So even though this is an obvious optimization for compilers, they will prioritize the serial guarantee over the optimization unless you tell it to relax that particular guarantee.
It's a mess and I agree with janwas: use a library (and in particular: use Google Highway) or something like Intel's ISPC when your hot path needs this.
- Aardwolf 3y agoHence my suggestion for language support. Add some language syntax where you can say: "add these doubles, order doesn't matter", which allows the compiler to use SIMD
- kolbe 3y ago-ffast-math if you want to push the onus on the compiler. Problem is you need to enforce that requirement on all user compilations, and I don't know what for MSVC. In-language would be nice.
- Aardwolf 3y ago-ffast-math is global though and can break other libraries. I mean local, vectorization compatible but portable syntax, without actually writing it (let compiler do the work)
- kolbe 3y agoTotally fair. I'm with you.
- eesmith 3y agoAccording to https://stackoverflow.com/questions/40699071/can-i-make-my-compiler-use-fast-math-on-a-per-function-basis https://stackoverflow.com/questions/40699071/can-i-make-my-c... you annotate it with: __attribute__((optimize("-ffast-math"))) I tried it with: double sum1(double arr[128]) { double tot = 0.0; for (int i=0; i<128; i++) { tot += arr[i]; } return tot; } __attribute__((optimize("-ffast-math"))) double sum2(double arr[128]) { double tot = 0.0; for (int i=0; i<128; i++) { tot += arr[i]; } return tot; } The Compiler Explorer (gcc 13.2, --std=c++20 -march=native -O3) generates two different bodies for those: sum1(double*): lea rax, [rdi+1024] vxorpd xmm0, xmm0, xmm0 .L2: vaddsd xmm0, xmm0, QWORD PTR [rdi] add rdi, 32 vaddsd xmm0, xmm0, QWORD PTR [rdi-24] vaddsd xmm0, xmm0, QWORD PTR [rdi-16] vaddsd xmm0, xmm0, QWORD PTR [rdi-8] cmp rax, rdi jne .L2 ret sum2(double*): lea rax, [rdi+1024] vxorpd xmm0, xmm0, xmm0 .L6: vaddpd ymm0, ymm0, YMMWORD PTR [rdi] add rdi, 32 cmp rax, rdi jne .L6 vextractf64x2 xmm1, ymm0, 0x1 vaddpd xmm1, xmm1, xmm0 vunpckhpd xmm0, xmm1, xmm1 vaddpd xmm0, xmm0, xmm1 vzeroupper ret It is still compiler-specific and non-portable, but at least it is not global.
- gpderetta 3y agoI'm guilty of having used it, but form the doc: " The optimize attribute should be used for debugging purposes only. It is not suitable in production code." It can be very unreliable.
- singhrac 3y agoOne of the things I struggled with in Rust’s portable_simd is that I think fadd_fast and simd intrinsics don’t mix well. This makes writing dot products only ok, rather than really easy. Here’s (one) reference: https://internals.rust-lang.org/t/suggestion-ffastmath-intrinsics-in-stable/14447/31 https://internals.rust-lang.org/t/suggestion-ffastmath-intri...