4 ms·
I cannot tell you if it is the solution they chose, but the fastest known implementation for single-precision FMA using double-precision multiplication and addi
by pascal_cuoq 8y ago
I cannot tell you if it is the solution they chose, but the fastest known implementation for single-precision FMA using double-precision multiplication and addition on conventional hardware is to use “round-to-odd” for the intermediate result. Any libc implementing “fmaf” for an architecture that doesn't have it has the same problem, and the solution looks like:
https://github.com/lattera/glibc/blob/b4d5b8b02133e0c317e6c836b51bbee3b00877b8/sysdeps/ieee754/dbl-64/s_fmaf.c https://github.com/lattera/glibc/blob/b4d5b8b02133e0c317e6c8...
“round-to-odd” is not a standardized rounding mode, even though it is so useful that some argue it should be. It can be emulated by the following sequence:
/* Reset rounding mode and test for inexact simultaneously. */
int j = libc_feupdateenv_test (&env, FE_INEXACT) != 0;
if ((u.ieee.mantissa1 & 1) == 0 && u.ieee.exponent != 0x7ff)
u.ieee.mantissa1 |= j;
The sequence has the drawback of not pipelining well on modern processors, but it's still faster than any known alternative, and at least the code is short.