4 ms·
That blog post is now a decade old, but includes an important quote: > The IEEE standard does guarantee some things. It guarantees more than the floating-point
by boulos 2y ago
That blog post is now a decade old, but includes an important quote:
> The IEEE standard does guarantee some things. It guarantees more than the floating-point-math-is-mystical crowd realizes, but less than some programmers might think.
Summarizing the blog post, it highlights a few things though some less clearly than I would like:
* x87 was wonky
* You need to ensure rounding modes, flush-to-zero, etc. are consistently set
* Some older processors don't have FMA
* Approximate instructions (mmsqrtps et al.) don't have a consistent spec
* Compilers may reassociate expressions
For small routines and self-written libraries, it's straightforward, if painful to ensure you avoid all of that.
Briefly mentioned in the blog post is IEEE-754 (2008) made the spec more explicit, and effectively assumed the death of x87. It's 2024 now, so you can definitely avoid x87. Similarly, FMA is part of the IEEE-754 2008 spec, and has been built into all modern processors since (Haswell and later on Intel).
There are still cross-architecture differences like 8-wide AVX2 vs 4-wide NEON that can trip you up, but if you are writing assembly or intrinsics or just C that inspect with Compiler Explorer or objdump, you can look at the output and say "Yep, that'll be consistent".
- nextaccountic 2y ago> There are still cross-architecture differences like 8-wide AVX2 vs 4-wide NEON that can trip you up Does those differences produce different results when running cross-platform simd code?
- boulos 2y agoDepends on what you're doing. The main issue here is reductions / accumulations. That is, if you have a bunch of floats like: float sum = 0.f; for (int i = 0; i < N; i++) { sum += x[i]; } and that you vectorize that to something like (typing in this comment, errors are likely): __mm256 sum_8wide = _mm256_setzero_ps(); for (int i = 0; i < N/8; i++) { sum_8wide = _mm256_add_ps(sum_8wide, _mm256_load_ps(x[8*i]); } // Now sum up the 8 values to get the final sum float sum = _mm256_hadd_ps(...); that will result in a different accumulation than if you did those are 4-wide and then a reduction. The usual solution is to either use the lowest common denominator (e.g., use AVX instead of AVX2) or more performance oriented, use the 4-wide SIMD units on ARM to "emulate" an 8-wide virtual vector (~15 years since I wrote NEON... and again, this is in a comment): float32x4_t sum_lo = ; // zero; float32x4_t sum_hi = ; // zero; for (int i = 0; i < N/8; i++) { sum_lo = vaddq_f32(sum_lo, vload1q_f32(x[8*i)); sum_hi = vaddq_f32(sum_hi, vload1q_f32(x[8*i + 4])); } // Reduce the sum in the same order You would want to write a "virtual SIMD" wrapper library, so you don't do this manually in lots of places.
- Someone 2y ago> but if you are writing assembly or intrinsics or just C that inspect with Compiler Explorer or objdump, you can look at the output and say "Yep, that'll be consistent". Surely people have written tooling for those checks for various CPUs? Also, is it that ‘simple’? Reading https://github.com/llvm/llvm-project/issues/62479 https://github.com/llvm/llvm-project/issues/62479, calculations that the compiler does and that only end up in the executable as constants can make results different between architectures or compilers (possibly even compiler runs, if it runs multi-threaded and constant folding order depends on timing, but I find it hard to imagine how exactly that would happen). So, you’d want to check constants in the code, too, but then, there’s no guarantee that compilers do the same constant folding. You can try to get more consistent by being really diligent in using constexpr, but that doesn’t guarantee that, either.
- boulos 2y agoThe same reasoning applies though. The compiler is just another program. Outside of doing constant folding on things that are unspec'ed or not required (like mmsqrtps and most transcendentals), you should get consistent results even between architectures. Of course the specific line linked to in that GH issue is showing that LLVM will attempt constant folding of various trig functions: https://github.com/llvm/llvm-project/blob/faa43a80f70c639f40021a175c60857bdcf7cc1f/llvm/lib/Analysis/ConstantFolding.cpp#L2407 https://github.com/llvm/llvm-project/blob/faa43a80f70c639f40... but the IEEE-754 spec does also recommend correctly rounded results for those (https://en.wikipedia.org/wiki/IEEE_754#Recommended_operations https://en.wikipedia.org/wiki/IEEE_754#Recommended_operation...). The majority of code I'm talking about though uses constants that are some long, explicit number, and doesn't do any math on them that would then be amenable to constant folding itself. That said, lines like: https://github.com/llvm/llvm-project/blob/faa43a80f70c639f40021a175c60857bdcf7cc1f/llvm/lib/Analysis/ConstantFolding.cpp#L1045 https://github.com/llvm/llvm-project/blob/faa43a80f70c639f40... are more worrying, since that may differ from what people expect dynamically (though the underlying stuff supports different denormal rules). Either way, thanks for highlighting this! Clearly the answer is to just use LLVM/clang regardless of backend :).
- mjcohen 2y agoYears ago, I was programming in Ada and ran across a case where the value of a constant in a program differed from the same constant being converted at runtime. Took a while to track that one down.