14 ms·
I don't understand all the crap that IEEE 754 gets. I appreciate that it may be surprising that 0.1 + 0.2 != 0.3 at first, or that many people are not educated
by twtw 8y ago
I don't understand all the crap that IEEE 754 gets. I appreciate that it may be surprising that 0.1 + 0.2 != 0.3 at first, or that many people are not educated about floating point, but I don't understand the people who "understand" floating point and continue to criticize it for the 0.1 + 0.2 "problem."
The fact is that IEEE 754 is an exceptionally good way to approximate the reals in computers with a minimum number of problems or surprises. People who don't appreciate this should try to do math in fixed point to gain some insight into how little you have to think about doing math in floating point.
This isn't to say there aren't issues with IEEE 754 - of course there are. Catastrophic cancellation and friends are not fun, and there are some criticisms to be made with how FP exceptions are usually exposed, but these are pretty small problems considering the problem is to fit the reals into 64/32/16 bits and have fast math.
- lenticular 8y agoYeah, the limitations of FP are well-known to anyone who does much numerical work. Floating point numbers are the optimal minimum message length method of representing reals with an improper Jeffery's prior distribution. A Jeffery's prior is a prior that is invariant under reparameterization, which is a mandatory property for approximating the reals. In this case, it is where Prob(log(|x|)) is proportional to a constant. Thus, we aren't going to ever do better than floats if we are programming on physical computers that exist in this universe. There is a reason why all numerical code uses them. Best to learn their limitations if you are going to use them, otherwise use arbitrary precision.
- erik_seaberg 8y agoThere's no reason for every step of a computation to be confined to the same very small message length. And the necessary error analysis should be built into the language, preferably in the same "advanced users only, here be dragons" package as the imprecise types themselves.
- TFortunato 8y agoSo interestingly, processor makers are on the same page with you re: computations, and lots of processors can internally do computations in "extended precision", e.g. 80-bit floats, only converting to/from 64-bit doubles at the start and end of the computation. https://en.m.wikipedia.org/wiki/Extended_precision https://en.m.wikipedia.org/wiki/Extended_precision
- garmaine 8y agoThat hasn’t been the case for over a decade.
- TFortunato 8y agoHow do you mean? The x86-64 instruction set / abi specifies long doubles as 80-bits, and still supports them ...
- leiroigh 8y agoAnd nobody uses this terrible mis-feature in practice, everything runs via 64 bit xmm registers. Rightly so, because programmers want their optimizing compiler to decide when to put a variable on the stack and when to elide a store/load cycle by keeping it in a register. With 80 bit precision, this makes a semantic difference and you end up in volatile hell.
- TFortunato 8y agoYeah I agree that everything typically runs in XMM registers and that's what people want. I'm not sure what about the availability of extended precision makes it s a misfeature? For some cases it IS what you want, and it's nice to be able to opt in to using it.. EDIT: If I had some application where I needed the extended range, like maybe I was going to run into the exact numbers above, I'd appreciate the ability to opt-in to this. Totally agree I wouldn't want the compiler to surprise me with it, but also not terrible, or useless. Code w/ Assembly: https://godbolt.org/z/W3ZmqJ https://godbolt.org/z/W3ZmqJ Output: https://onlinegdb.com/Sy_I3Q1ME https://onlinegdb.com/Sy_I3Q1ME
- 3pt14159 8y agoOutside of the academic world decimals are almost always a better solution if performance isn't critical. Most logic is multiplicative. For example, apply a 30% tax on a dollar quantity and display both subtotal and grand total. With floats, there are inequalities. With decimal there usually aren't unless you're dividing, but we already have to deal with divide errors in base ten, and it is much more likely to need to represent 0.30 than 1/3 and because decimal shares a base with binary (since it's factors are 5 and 2) binary doesn't really get us anything but headaches anyway. It's true that there are still gotchas, but they happen less often and usually don't end up looking stupid and weird for no reason. That 0.1 + 0.2 = 0.300000000000001 is dumb and we all know it.
- ChrisLomont 8y agoYou cannot do anything in finance with such reasoning. Take something simple, say a mortgage at 5% compounded 12 times a year. To compute payments using some fixed length representation or decimal is going to lead to more error than to use the usual floating point. This rabbit hole would continue for many applications. Floating point makes them all much easier to do well.
- jes5199 8y agoHave you worked on finance software? I have - we always used ints for everything, so we could avoid rounding suprises
- bigiain 8y agoI did a web project in the gambling space ~10 years back - we were legally required to perform all calculations as integers in ten thousandths of a cent (or millionths of a dollar). We chose to _not_ do _any_ calculations client side in Javascript...
- dotancohen 8y agoWhich regulation is that? I've worked on financial applications, but not gambling, and I've not heard of this regulation. I should probably know about it!
- vanderZwan 8y ago> Thus, we aren't going to ever do better than floats if we are programming on physical computers that exist in this universe. Maybe, but that doesn't mean the particular implementation of floats being used is the best one. See also: Unums and Posits https://posithub.org/ https://posithub.org/
- nestorD 8y agoTo my understanding, experiments with Unums showed shortcoming that Gustafson didn't anticipate and lead to Posits which drop the fixed length constraint. Doing that makes improving precision a lot easier but at the cost of computation time. Overall I am not convinced that the current implementation is optimal but it is a very good trade-off between speed and precision.
- vanderZwan 8y ago> Doing that makes improving precision a lot easier but at the cost of computation time. Not quite. The difference in computation time is the current the lack of hardware support, not something inherent to the underlying encoding method. So in practice you are right, but in, for example, embedded contexts without floating point hardware, the performance advantages of IEEE floats should disappear (especially if using a 16 or 8 bit posit suffices). Posits are simpler to implement than IEEE floats (less edge cases) and use more bits for actual numbers whereas IEEE floats waste about half on NaNs. The use of tapered precision is also nice.
- tntn 8y agoEven if hardware support existed, it seems like a variable length encoding has some inherent overhead relative to a fixed length encoding. If you have a "base length" of e.g. 32 bits and occasionally expand to 64, there's an inherent cost there in both computation and memory, presumably for greater precision. Perhaps that overhead could be minimal with hardware support, but it seems it must have some.
- 8y ago
- erpellan 8y agoThe limitations should be well known. One of the first things I check when joining a financial software project is how the system represents money. I’m rarely surprised. (It’s inevitably floats or doubles)
- ayidnelm 8y ago> Floating point numbers are the optimal minimum message length method of representing reals with an improper Jeffery's prior distribution. Do you have a link to a proof or discussion of this? I haven't heard this before and I would love to have this statement unpacked a little more.
- chrisseaton 8y agoPeople get upset that floating point can’t represent all infinite number of real numbers exactly - I can’t understand how they think that’s going to be possible in a finite 64 bits.
- baddox 8y agoOr on any computer at all, even an “infinite” (at least unbounded) computer like a Turing machine, considering that almost all real numbers are not computable.
- m0zg 8y agoTo hit the point home a little harder: you can easily iterate through the entire representable set of float32 on a modern machine within seconds. I've encountered many engineers who don't quite get that.
- chrisseaton 8y agoRight - if you have a monadic function that takes a 32-bit float, your tests should probably literally cover every single input value.
- jancsika 8y agoWait, where did OP's 64-bit slot go? > I can’t understand how they think that’s going to be possible in a finite 64 bits. You apparently stole 32 of them to make your bat. If you put them back your tests balloon to half a century each.
- chrisseaton 8y agoWow that's aggressively snarky. I presume they were saying 'and for 32-bit floats you also get this property that you can...'
- tom_ 8y agoAn alternative calculation: https://news.ycombinator.com/item?id=18109432 https://news.ycombinator.com/item?id=18109432 "You can rent a Skylake chip on Google Cloud that'll perform 1.6 trillion 64 bit operations per second for $0.96/hr preemptively. That's enough to run one instruction over a 64 bit address space exhaustively over 120 days, or for ~$2800" It might not make economic sence to actually make this happen for any realistic test, but it's interesting that it might actually be feasible to do it on any kind of human timescale...
- Gibbon1 8y ago> exceptionally good way to approximate You answered your question. 99% of the time being exact is a requirement and calculation speed is utterly unimportant, thus using IEEE 754 results in programs that are fundamentally broken.
- coddingtonbear 8y agoIs that really true? In my experience, 99.9% of the time I don't need an exact number; the vanishingly few times when I have such a need (almost entirely calculations involving currency), using a fixed point representation is simple enough.
- Gibbon1 8y agoYou don't need an exact number but customer data is universally decimal. Soon as you blindly convert that to IEE754 everything is now broken.
- coddingtonbear 8y agoIs it really, though? I'm honestly struggling to think of a non-currency situation in which fractional customer data necessarily be handled as a decimal value -- and, honestly, even if the availability heuristic might make them seem more common than they are, I'd be astonished if even a single percent of general calculations programmers collectively ask computers to perform are involving currency. Most real-life situations just don't even inherently _have_ that kind of precision, let alone need it. Seriously, I can't think of a time when I've needed to store a coordinate or a person's height as a decimal value to prevent something from being broken.
- tedunangst 8y agoMy 5/8s wrench disagrees. Happens to store quite nicely in a float, however.
- dkarl 8y agoPeople do different kinds of work, so there are programmers who experience it both ways, 99% of the time floats are good solution or 99% of the time floats are an incorrect solution. Because of history and language support, classes and other resources for learning to program teach you to use floating-point numbers and don't bother with alternatives. As a result you have a lot of programmers who default to treating every number with a dot in it as floating point number, and they get burned by it, and instead of realizing it's just a gap in their education that they can correct, they treat overuse of floats as a mistaken industry-wide consensus that needs to be overturned.
- svat 8y ago> considering the problem is to fit the reals into 64/32/16 bits and have fast math Floating-point numbers (and IEEE-754 in particular) are a good solution to this problem, but is it the right problem? I think the "minimum of surprises" part isn't true. Many programmers develop incorrect mental models when starting to program, and get no feedback to correct them until much later (when they get surprised). It is true that for the problem you mentioned, IEEE 754 is a good tradeoff (though Gustafson has some interesting ideas with “unums”: https://web.stanford.edu/class/ee380/Abstracts/170201-slides.pdf https://web.stanford.edu/class/ee380/Abstracts/170201-slides... / http://johngustafson.net/unums.html http://johngustafson.net/unums.html / https://en.wikipedia.org/w/index.php?title=Unum_(number_format)&oldid=873507492 https://en.wikipedia.org/w/index.php?title=Unum_(number_form... ). But many programmers do not realize how they are approximating, and the "fixed number of bits" may not be a strict requirement in many cases. (For example, languages that have arbitrary precision integers by default don't seem to suffer for it overall, relative to those that have 32-bit or 64-bit integers.) Even without moving away from the IEEE-754 standard, there are ways languages could be designed to minimize surprises. A couple of crazy ideas: Imagine if typing the literal 0.1 into a program gave an error or warning saying it cannot be represented exactly and has been approximated to 0.100000000000000005551, and one had to type "~0.1" or "nearest(0.1)" or add something at the top of the program to suppress such errors/warnings. At a very slight cost, one gives more feedback to the user to either fix their mental model or switch to a more appropriate type for their application. Similarly if the default print/to-string on a float showed ranges (e.g. printing the single-precision float corresponding to 0.1, namely 0.100000001490116119385, would show "between 0.09999999776482582 and 0.10000000521540642" or whatever) and one had to do an extra step or add something to the top of the program to get the shortest approximation ("0.1").
- crankylinuxuser 8y agoWhen NASA can't even get it right, because of "surprises", there's no chance in hell I'm blaming us mere mortal programmers... or even 10x wizards. (0) It's time to look at other ways to depict fractional parts of numbers in a computer. I know that one can express any rational number as a integer fraction. And our computers are incapable of expressing a irrational number exactly - it does so to a certain precision... In other words, every number a computer expresses is a rational number. The exception is if the computer could express irrational numbers as symbolics, then we could work with the symbolic instead. And then as a last pass, the symbolic could convert to a imprecise rational depiction, or express as its native type. (0) https://itsfoss.com/a-floating-point-error-that-caused-a-damage-worth-half-a-billion/ https://itsfoss.com/a-floating-point-error-that-caused-a-dam...
- ummonk 8y agoPresumably we could actually make decimal floating point computation the default and greatly reduce the amount of surprise. I don't think the performance difference would be an issue for most software.
- smallnamespace 8y agoDecimal floating point won't avoid this issue, for a sufficiently large value the ulp would be 10.
- ummonk 8y agoIt would solve more common issues like this though: > I appreciate that it may be surprising that 0.1 + 0.2 != 0.3 at first, or that many people are not educated about floating point, but I don't understand the people who "understand" floating point and continue to criticize it for the 0.1 + 0.2 "problem." That's not a calculation that should require a high level of precision.
- smallnamespace 8y agoThe correct solution is to understand how floating point number systems work and use near comparisons for floats. Decimal fp is still 'wrong' for, say, 1/3 + 1/3 = 2/3.
- int_19h 8y agoA lot of real-world data is already in base-10 for obvious reasons, and so an arrangement that lets you add, subtract and multiply those without worrying is worthwhile, even if it can't handle something more exotic.
- smallnamespace 8y agoWould you really call 'any rational with divisible factors other than 2 and 5' to be 'exotic'? Maybe we really should move back to base-60 like the Babylonians used, then you could at least divide by 3.
- etCeteraaa 8y agoBecause if there are obvious edge and corner cases, like overflow scenarios, a professional system will either ensure that expectations are lived up to, or flatly denied as errors. No surprises.
- delhanty 8y agoExactly! Very far from a floating point expert here, but what I do is to scale-down by a few odd prime-power factors as appropriate: Scaling down by powers of 5 is obviously appropriate for decimals, currency etc. Scaling down by powers of 3 is good for angles measured in the degrees, minutes, seconds system. If one scales down a lot there is an increased risk of overflow, so one can compensate by scaling up some powers of 2. The way I think of this is as using my own manual exponent bias [0]. >the exponent is stored in the range 1 .. 254 (0 and 255 have special meanings), and is interpreted by subtracting the bias for an 8-bit exponent (127) to get an exponent value in the range −126 .. +127. So, for example, even single-precision number are always exact multiples of 1/(2^126), and I'm just changing the denominator to contain powers of 3, 5, 7, ... etc. [0] https://en.wikipedia.org/wiki/Exponent_bias https://en.wikipedia.org/wiki/Exponent_bias
- microcolonel 8y agoIntegers are a lot less trouble for many currency problems, but I think some people are afraid of multiplying integer fractions. In financial calculations I've seen, figures are given in standard magnitudes (per cent, per mille, basis points, integer cents, etc.) which, if you're lucky with your language, can be encoded as types which can be promoted to higher precision (somewhat) transparently.
- hedora 8y agoI took the table to be a handy guide to where arbitrary precision is the default vs. hw accelerated math. Filtered by languages I care about, I guess I have no choice but to learn perl 6 if I want correct (but presumably slow) floating point with elegant syntax (my taste might not match yours). I’d be curious to know what the random GPU languages and new vector instruction sets do with this computation. I don’t think they’re all 754 compliant.
- twtw 8y agoCan't comment on the situation with other GPU languages, but CUDA on GPUs since fermi are 754 compliant, with the exception that certain status flags are unavailable.
- yoz-y 8y agoTo me the only downside of IEEE 754 is that most languages including C and C++ do not provide a sensible canonical comparison methods. This leads to surprised beginners and then a ton of home made solutions which are often not appropriate.
- MauranKilom 8y agoWould you really want a default comparison where a == b does not imply a - b == 0?
- yoz-y 8y agoI think it depends, in languages which have implicit type coercion I think that would hurt. In languages like swift, where you need to explicitly cast even an Int to Double it would be less of a footgun. I'd rather floats have some overloaded operator maybe ~=, for approximate comparison.