6 ms·
Perspectives on Floating Point
- cosignal 2y agoVery nice graphics in this.
- rhythane 2y agoA really interesting review. The idea of relative error makes sense in most cases, but when we need to do subtraction and difference matters, maybe absolute error is actually better.
- exmadscientist 2y agoIf you are doing only additive operations, then, yes, absolute error might actually be the best choice. But as soon as multiplications start to show up, they are enough trouble that they tend to dominate the whole error propagation show. Since many real calculations have multiplication in them, you end up having to optimize the whole thing for multiplicative operations, and so we end up just using relative errors everywhere. You can, of course, do a very specialized optimization for one particular algorithm, but that tends to not be a very good use of time. Usually. (Counterexample: Kahan summation!)
- dahart 2y agoSubtraction is the very case where relative error might matter most, and error relative to magnitude of the original unsubtracted numbers causes the most surprises. Absolute error has useful applications, without any doubt, and regardless of arithmetic operation, but it probably doesn’t make sense to say it’s “better” without a specific problem in front of us, and without specific goals and priorities. Error tolerance is always up to the user.
- yxhuvud 2y agoOne thing I think would be nice for floating point numbers, is that I'd prefer if there were two separate types - one where NaN and the two infinites are allowed, and one where they are not allowed but instead emit an error. The former would be used by some few mathematicians etc, and the rest of us could use the latter. The upside would be better error handling close to the source of the issue, and better optimizations as the not-normal values throw a wrench into optimizing math.
- hollerith 2y ago>better optimizations as the not-normal values throw a wrench into optimizing math. I'm not an expert on architecture, but I would've guessed that the need to branch to process the error would be the wrench and that the use of NaN and the infinities allow better optimizations.
- exmadscientist 2y agoIEEE-754 NaN and Infinity have nothing to do with optimization. They come straight from math: * What is +1/0? It has to be +Infinity -- nothing else will do. * What is -1/0? It has to be -Infinity -- nothing else will do. * What is 0/0? There's no way to tell from the information we've got -- it's undefined: Not A Number. (However, should 0/0 come up as a result of taking the quotient of two functions that happen to both reach zero at a point, then sometimes the limit of that quotient is meaningful, and might have a numerical result.) IEEE-754 chose to signal these things in-band, so we get NaN and Infinity to deal with in our floats and doubles.
- hollerith 2y agoNaN and Infinity have nothing to do with optimization only if your mental model of CPUs is extremely simplistic.
- fweimer 2y agoI'm pretty sure IEEE 754 covers trapping floating point math. It's not just about in-band signaling. It's unfortunate that the standard is proprietary, so we can't easily reference it to figure out what it says. Anyway, most CPUs support a trapping mode, after all. Here's an example with glibc: #define _GNU_SOURCE #include <fenv.h> volatile double x = 1.0; volatile double y; volatile double quotient; int main(void) { feenableexcept(FE_DIVBYZERO); quotient = x / y; } As far as I understand it, overall support for trapping math is poor because not much code is trapping-aware. It would be quite annoying if JSON parsing results in SIGFPE due to an Inexect trap, for example.
- 2y ago
- kolbusa 2y agoNot sure why the article does not reference the following paper which is a must read for anyone working with floating point: https://docs.oracle.com/cd/E19957-01/806-3568/ncg_goldberg.html https://docs.oracle.com/cd/E19957-01/806-3568/ncg_goldberg.h... (original: https://www.itu.dk/~sestoft/bachelor/IEEE754_article.pdf https://www.itu.dk/~sestoft/bachelor/IEEE754_article.pdf).
- dunham 2y agoI also recently saw this paper on the difficulty of solving the quadratic equation with floating point numbers: https://cnrs.hal.science/hal-04116310/document https://cnrs.hal.science/hal-04116310/document And also Gerald Sussman saying: > The only thing that scares me in programming is floating point. https://youtu.be/Tdwr9tweTDE?t=1145 https://youtu.be/Tdwr9tweTDE?t=1145
- andrepd 2y agoPosits https://posithub.org/docs/Posits4.pdf https://posithub.org/docs/Posits4.pdf are an excellent perspective for an alternative to IEEE floats.