3 ms·
My understanding is that in this case it's GCC's optimizer evaluating the expression at compile time in a way that gives a different result to if it's done at r
by dpwm 6y ago
My understanding is that in this case it's GCC's optimizer evaluating the expression at compile time in a way that gives a different result to if it's done at runtime.
EDIT: GCC gets it wrong with -O0 (when it's evaluated at runtime) and right with -O2 (when it evaluates it at compile time)
Clang appears to get it right when evaluated at compile time or runtime.
Clang uses the divsd instruction; GCC uses the fdiv instruction – so this really is SSE vs x87 FPU.
- nwallin 6y agoclang gets it "right" because it uses the sse divsd instruction to perform the division, gcc gets it "wrong" because it uses the x87 fdiv instruction. You can configure gcc to use sse math with the -mfpmath=sse instruction. The author states that gcc still gets it "wrong" but he simply is not correct. gcc will use divsd and gets 0.501782303180000055. gcc defaults to use the x87 fpu instead of sse because not all 32 bit x86 cpus have SSE instructions. It's a safer default. When compiled in 64 bit mode, it uses up to SSE2 instructions, because all x86_64 CPUs have SSE2. https://godbolt.org/z/3PfKQU https://godbolt.org/z/3PfKQU
- ndesaulniers 6y agoI think `-mno-see -mno-sse2` are of interest here, too.