3 ms·
double-double or quad precision is necessary to emulate double-precision FMA. The article is talking about emulating single-precision FMA with double-precision.
by pascal_cuoq 8y ago
double-double or quad precision is necessary to emulate double-precision FMA. The article is talking about emulating single-precision FMA with double-precision. The relevant paragraph is:
> Luckily the vast majority of floating-point math in games is done to float (32-bit) precision, and I was quite happy to use double (64-bit precision) instructions in the emulation of FMA.
One big difference between quad-precision and double-double is that quad-precision has a much wider exponent range. If you use double-double, you need to worry that the result of the multiplication may underflow a double. In the article you cited:
> First, the value ul has to be the error term of the multiplication a · b, in order to avoid some degenerate underflow cases: the error term becomes so small that its exponent falls outside the admitted range. […] The algorithm will behave correctly even if some computed values are not normal numbers, as long as ul is representable.
… implying that the algorithm may not compute the FMA if the error of the multiplication is not representable, which can happen when it is below the normal range.
- lifthrasiir 8y agoProbably I was not as careful in the choice of a word "double-double". My point was that the error recovery (or, as the paper refers, the error-free transformation) seems crucial for FMA emulation in general. You are entirely right that double-double has many pitfalls. My thinking was that, in this particular case we need around 24 × 3 = 72 bits of mantissa (I haven't verified the exact number, but it clearly exceeds 60 bits) to avoid the double rounding---which double precision cannot provide. The verified algorithm gives a lot more than enough headroom for this particular setting: ExactMult is just a normal double multiply and ExactAdd will recover the error out of double addition. It might even be possible to optimize later cases. But it seems to me that you can't really get rid of the error recovery procedure itself. Well, I may be wrong. EDIT: Oh, I see your neighboring replies. So I was wrong! The glibc solution however looks pretty expensive and it is unfortunate that there exists no faster alternatives known.
- pascal_cuoq 8y agoThere is no reason to estimate the required precision as 3 times the original precision, because floating-point addition does not work like that. If you want to compute the exact result of a floating-point addition, you need approximately emax - emin bits of precision. Floating-point addition is never computed this way. On the other hand, multiplication does have the property that the required precision for representing the result of multiplying numbers with precisions p and q is p+q.
- lifthrasiir 8y agoWe don't compute the exact sum, we just need enough precision to ignore the double rounding. The most pathological cases are therefore either: - the product is just above the ULP of the addend, or - the addend is just above half the ULP of the product. I'm not sure about the latter (the possible bit patterns of the product are constrained) but the former clearly requires 3 times the original precision, and beyond that there is no possibility of double rounding. The same thing can be said for the latter. Of course all these points are moot when it is known that rounding-to-odd can be used to avoid error recovery at all.