12 ms·
Weird compiler bug – Same code, different results
- brucehoult 6y agoIf you read the article you find that this is a bug in MinGW64 libraries. It's not a compiler bug, it's not a problem in C/C++, or even in IEEE-754 floating point math. The MinGW64 thread library simply forgot to initialize the FPU correctly. He'd have had the same problem if he wrote his code in assembly language and then ran it in a MinGW64 thread.
- zaitanz 6y agoYes this is technically true, but it manifests itself through use of the MinGW compiler. When I first noticed the problem, there was very little indication as to the cause. TBH I accredit dumb luck more than anything to finding the cause of this. It took many hours.
- gus_massa 6y ago> This means that floating point arithmetic is non-associative. In that A + B != B + A. The equation is the commutative property, not associative. IIRC addition in IEEE-754 is commutative. The property that fails is (A+B)+C = A+(B+C).
- zaitanz 6y agoOp Here. Yes you are correct. Been very sleep derived when I wrote this (newborns eh). I'll update this. I didn't want to go too far into the background of floating point math in the post as it wasn't ultimately the issue, just something I was mindful of.
- ojnabieoot 6y agoI think you made the wrong correction in the article: > Update: Thanks to gus_massa and wiml @ HackerNews for pointing out I had used associative instead of commutative.... This means that floating point arithmetic is noncommutative [i]n that A + B != B + A. Floating-point arithmetic is commutative, but not associative. So for floating-point numbers, A + B = B + A always, but A + (B + C) != (A + B) + C in many cases. The problem comes about from significant digits and rounding - A + B might round up to 1, but B + C might round down to zero, etc. The point remains that the order in which you do things actually does matter in floating-point arithmetic, but it's the order you compute the "inner parentheses," not the actual order of numbers in (e.g.) the FPU registers.
- zaitanz 6y agoThank you for your reply. I'm going to just lift your explanation as it's much better than my own.
- titzer 6y agoDo not use C/C++ for numerical code where accuracy is needed. These languages are not specified to conform to IEEE 754 and you are absolutely asking for trouble.
- cperciva 6y agoMy copy of C11 says Annex F (normative) IEC 60559 floating-point arithmetic F.1 Introduction This annex specifies C language support for the IEC 60559 floating-point standard [...] previously designated ANSI/IEEE 754-1985.
- titzer 6y agoC "supports" whatever the hardware and compiler feel like supporting. It absolutely does not mandate anything. In particular, there are a number of particularly simple compiler optimizations that are not forbidden, though they are not technically correct according to IEEE 754, such as algebraic reassociation and simple commutativity. Moreover, C allows subexpressions to be computed in higher precision (e.g. 80 bit "long double"), which is observable. That last one is primarily due to the x87 FPU coprocessor design that has given us a good 35 years of headaches. Good riddance to that!
- shakna 6y ago> Support for Annex F (IEEE-754 / IEC 559) of C99/C11 > The Clang compiler does not support IEC 559 math functionality. Clang also does not control and honor the definition of __STDC_IEC_559__ macro. Under specific options such as -Ofast and -ffast-math, the compiler will enable a range of optimizations that provide faster mathematical operations that may not conform to the IEEE-754 specifications. The macro __STDC_IEC_559__ value may be defined but ignored when these faster optimizations are enabled. If you make use of Clang, you won't find support for Annex F, and in point of fact you can't even macro check to see if it even will support Annex F. The story with GCC is... More complicated. It guarantees it'll follow Annex F for only some operations [0]. So the upshot is... Most people can't tell when and how Annex F might be followed. [0] https://gcc.gnu.org/onlinedocs/gcc-7.4.0/gcc/Floating-point-implementation.html https://gcc.gnu.org/onlinedocs/gcc-7.4.0/gcc/Floating-point-...
- lysium 6y agoI would have never thought that the output of some floating point operations depend on the „reset of the floating point package“. What gives?
- willxinc 6y agoThings like rounding mode and denornal behavior could be controlled by CPU flags. Imy not familiar with x86, but this is defiy the case in PPC.
- lysium 6y agoYes, that makes sense, thank you!
- cperciva 6y agoProbably the author is accidentally getting 80-bit "extended double" arithmetic.
- cozzyd 6y agoThat should be clear from looking at the assembly, shouldn't it? Don't think the floating point environment can change that but I could be wrong. Floating-point rounding modes seem more likely to me. The author should be able to dump the floating point configuration to confirm, I'm sure.
- cperciva 6y agoThe same x86 instructions are used for arithmetic on float/double/extended types; the difference is a precision setting in the x87 control word.
- cozzyd 6y agoWell in most cases I've looked at generated assembly (not that often), the xmm registers are used even for scalar operations, which I thought was the default option for gcc on x86-64, but I suppose it might differ on different systems (or perhaps 32-bit mode was used for some reason).
- deleted 6y ago[deleted]
- peter_d_sherman 6y ago>"Compiler bug? But how… Looking at the code, the only things that could create the bug were the operators and pow() call. These are part of the C++ standard offering so it’s highly unlikely that either of these would have an error. What next? Off to GodBolt to try some other compilers to see if I get wonky results on them. MSVC++… nope, Clang… nope, GCC-Trunk on Linux… nope. I am only getting this weird behaviour in TDM-GCC. Time to try some other MingW64 compilers. Nuwen… Yes, Mingw64.. yes, TDM-GCC… yes. Confirmed: It’s a bug in MinGW64. When I create a new thread and run floating point operations in that thread, I get slightly different answers." Interesting!