11 ms·
Floating point numbers, and why they suck
- RicoElectrico 5y agoI wish that fixed point numbers (as in Qx.y) was a first class citizen in programming languages. As well as saturation arithmetic. Actually I wonder if anybody did a performance/power comparison of using floating point vs fixed point math in some common tasks using modern CPUs with extensive FP support.
- okl 5y agoThey are a first class citizen in Ada, you even get two (three) varieties: http://www.ada-auth.org/standards/12rm/html/RM-3-5-9.html http://www.ada-auth.org/standards/12rm/html/RM-3-5-9.html / http://www.ada-auth.org/standards/12rm/html/RM-3-5-7.html http://www.ada-auth.org/standards/12rm/html/RM-3-5-7.html
- adgjlsfhk1 5y agoJulia has this https://github.com/JuliaMath/FixedPointNumbers.jl https://github.com/JuliaMath/FixedPointNumbers.jl. In general, fixed point is a little slower, but it's not awful. Fixed point is really good for small bit-widths in hardware though.
- togaen 5y agoFloating point behaviors are well understood. This type of problem shouldn’t happen if your developers are qualified.
- m4r35n357 5y agoAbsolutely correct. Poor education.
- h2odragon 5y agoThis is an issue that has bitten people before. Indeed, there's probably been some discussion of it and might even be some common mitigations taught as "best practice" in some places. Do we ever need to store floats? using them in flight is one thing, but stored data so often needs to be some fixed precision.
- okl 5y agoIf you want to type pun floats, store hexadecimal literals. Those will be accurate.
- kstenerud 5y agoThe problem is that we're still stuck with only binary floating point types in our CPUs and compilers and runtime environments, over a decade after ieee754 released their decimal float spec [1]. Once we finally move away from binary floats, these nasty binary-exponent-decimal-exponent lossy conversions will end (as will the horribly complicated nasty scanf and printf algorithms). This is a solved problem. Our actual problem now is momentum of adoption. I've anticipated that this will eventually happen (hopefully sooner rather than later), and proactively marked binary floats as "legacy formats" [2]. [1] https://en.wikipedia.org/wiki/Decimal64_floating-point_format https://en.wikipedia.org/wiki/Decimal64_floating-point_forma... [2] https://github.com/kstenerud/concise-encoding/blob/master/ce-structure.md#binary-floating-point https://github.com/kstenerud/concise-encoding/blob/master/ce...
- apaprocki 5y agoMany compilers and runtimes support them just fine. CPUs don’t really support them, with the exception of POWER. The dirty little secret (use case measurement), however, is that Intel’s software library on x86-64 is faster than POWER’s instructions. I don’t believe there is any roadmap (e.g., 5 years out) to add hardware support mainly because customers don’t need them or want them (globally speaking) and software is fast enough for majority of use cases.
- vintermann 5y agoI agree that's cool. It's usually sold as for finance applications, for compliance with rounding rules from the pre-binary era etc. But really, any number that lived as a decimal digit string somewhere in its lifetime should probably have been decimal throughout. But I don't think it would have mattered here anyway. With finite precision floating point addition is not going to be associative, decimal or not.
- kstenerud 5y agoIn practice they don't need to be perfectly associative across all possible values, because most real-world calculations tend to stay within a few orders of magnitude, and don't require more than 10-15 significant digits (even for Earth orbital calculations). Once you need more than that, you'd have to use a multi-word numeric format regardless (same as if your operations started overflowing your largest int type).
- ape4 5y agoAncient COBOL had binary coded decimal.
- marcosdumay 5y agoEvery database and most programing languages have decimal floating point types. They usually don't use BCD, and instead use some more efficient representation. But anyway, the representation is irrelevant. By the way, integral BCD (what COBOL did most of the time) is useless.
- abujazar 5y agoFloating point arithmetic doesn't «suck», it's the code that treats floats as decimal numbers that sucks.
- okl 5y agoI'm very annoyed by this title. It's a beginners mistake and FP numbers just are what they are -- a tradeoff between precision, range and efficiency. If you use FP, you should be familiar with its intricacies in the same way a C programmer needs to be aware of unsigned overflows.
- ozychhi 5y agoNooo unsigned integers suck, if you subtract 5 from 3 what would you get? Then I had a stroke of genius and figured it out, I have no idea what I'm doing.
- pjam 5y agoThis. Sure, floating point numbers come with a footgun, but the title is so clickbaity, would it really hurt to name it: "What I wish I knew about floating numbers before relying on them?"
- spicybright 5y agoI think it's common to just not be taught floating points need to be handled a lot differently than whole numbers, so the common experience is getting bit by it. I don't remember ever being taught that in my traditional schooling, personally. This is a good link to send to people: https://floating-point-gui.de/ https://floating-point-gui.de/
- maxlybbert 5y agoI remember covering how floating point were stored, but I don’t remember going all that deep into how that affected arithmetic.
- longemen3000 5y agoalso: https://0.30000000000000004.com/ https://0.30000000000000004.com/
- Mikhail_K 5y agoIEEE 754 floating-point format has been very carefully and deliberately designed by the top experts in the field. When someone "it sucks", that is a guarantee that whatever design he might come up with would suck ten times as much.
- aix1 5y agoObligatory: What Every Computer Scientist Should Know About Floating-Point Arithmetic https://docs.oracle.com/cd/E19957-01/806-3568/ncg_goldberg.html https://docs.oracle.com/cd/E19957-01/806-3568/ncg_goldberg.h...
- Jensson 5y agoEveryone knows this from using calculators in school, even non programmers knows this.
- mgaunard 5y agoAsking basic questions about floating-point in interviews remains to this day a good way to distinguish good and bad programmers. You'd think everyone would know the basics, but even in stuff like finance, people routinely fuck up decimals.
- longemen3000 5y agoIn Julia, there is the `isapprox` function to inexactly compare 2 numbers. (you can use it in infix form: `x≈y`. By default, two numbers are approximately equal if their relative tolerance is less than `sprt(eps(typeof(x))` (around 1e-8 for 64bit Floating point numbers) Using equality to compare 2 floating point numbers doesn't get you anywhere, specially if you are using mathematical functions (the implementation of `sin` or `exp` could vary with operating systems or software versions)
- okl 5y agoRelative tolerance becomes problematic though if your numbers are very close to zero. (https://randomascii.wordpress.com/2012/02/25/comparing-floating-point-numbers-2012-edition/ https://randomascii.wordpress.com/2012/02/25/comparing-float...)
- adgjlsfhk1 5y agoWhich is why it also lets you specify an absolute tolerance. That's not done by default because there isn't a good way to choose one (unlike with relative tolerance).
- longemen3000 5y agoyes, it is a common problem, it is acknowledged in the function's documentation: https://docs.julialang.org/en/v1/base/math/#Base.isapprox https://docs.julialang.org/en/v1/base/math/#Base.isapprox
- fiedzia 5y ago> We are sending the data to the warehouse as a float There is your problem
- unnouinceput 5y agoAnd that's why you user "currency" type of style. Both as calculation in your programming language (Go in this article) and in DB (PostgreSQL in this article). Fixed width "floating" number which is actually your largest integer type (64 bits these days, but there are libraries for 128 bits aplenty as well) is your friend.
- nomorecommas 5y agoYep, floats should be considered dangerous. $ python3 -c 'print(.1 + .2)' 0.30000000000000004
- f154hfds 5y agoAny vector addition done concurrently will result in indeterminate roundoff errors (to be precise there might be O(N!) different possible outputs depending on the order of calculation), this is literally the point of the field of numerical analysis. But that's really the tip of the iceberg. Beware any mathematical operation done with numbers of vastly different orders of magnitude. Beware spinning your own matmul even if you can't guarantee some normalcy to the numbers in your matrices. Calculating an average is fairly painless, you can create an upper bound on the possible error that isn't usually dangerous. This isn't the case for all computations.
- pdpi 5y agoThe vast majority of problems with floating point numbers are simply due to people not understanding number bases. Only a tiny subset is caused by people not understanding floats proper. Consider the number 1/3. In ternary (base 3), you write that as 0.1, whereas in decimal (base 10) you write it as 0.3333... recurring. If you try to represent that number with a fixed number of decimal places, you have precision issues. E.g. decimal 0.333 converts to 0.02222220210 in ternary. Now, the thing with that example is that we treat decimal as a special, privileged representation, so we accept that 1/3 doesn't have a finite representation in decimal as an entirely natural fact of life, while we treat that same problem converting between decimal and binary as a fundamental deficiency of binary. Let's talk about why some numbers have finite representations while others don't. 53/100 is written as 0.53 in decimal. The general rule is that if the denominator is a power of ten, you just write the numerator, then put the decimal mark have as many digits from the right as the power of ten. If the denominator is not a power of ten, you make it one. 1/2 turns into 5/10, which is 5 with one decimal place, or 0.5. Obviously, you can't actually do this for 1/3. There's no integer `a` where 3a is a power of ten. The general rule here is that, if the denominator has any prime factors not present in the base, you don't have a finite representation. 4/25 = 16/100 = 0.16 has a finite representation, but 5/7, 13/3 don't. Now, because 3 is coprime with 10, no number with a finite ternary representation can have a finite decimal representation (and vice versa), and we're used to it. Where binary vs decimal becomes tricky and confuses people is that 10 = 2 * 5, so numbers that can be expressed with a power-of-two denominator have finite representations in both binary and decimal, so you can convert some numbers back and forth with no loss of precision. Numbers with a factor of 5 somewhere in their denominator can have finite decimal represntations (1/5 = 2/10 = 0.2), but can't have finite representations in binary. 0.1 = 1/10 = 1/(2 * 5) and you can't get rid of that five. And that's why everybody gets bitten by 0.1 seeming to be broken.