6 ms·
A precise integer value is only guaranteed to be representable losslessly in a double if it is up to `64 - 1 (sign) - 11 (exponent) = 52` bits in magnitude. Th
by ThreeFx 7y ago
A precise integer value is only guaranteed to be representable losslessly in a double if it is up to `64 - 1 (sign) - 11 (exponent) = 52` bits in magnitude.
This should be fairly obvious with knowledge about how floating point numbers are represented internally IMO.
Edit: Be more precise about what can be represented.
- yoz-y 7y agoI've seen quite a lot of errors stemming up from assumptions about floating point numbers. Not sure there is a good way of handling this in the end, except exert caution. Even basic assumptions like f + 1 > f will have this issue.
- magicalhippo 7y agoThe problem is that in most languages, to novices floating point numbers swim like a duck, quack like a duck, but it turns out they're alligators. A good way to get around this is reading https://floating-point-gui.de/ https://floating-point-gui.de/ to weed out any preconceptions, but yeah it's difficult to steer novices there without them stepping on one of the pitfalls first.
- stephencanon 7y agoNote that f + 1 > f can also fail for integers in many languages (e.g. it does not hold for signed integers in C or C++, because the behavior is undefined when you add 1 to INT_MAX, and unsigned integers always wrap around). This particular gotcha is not unique to floating-point.
- nayuki 7y agoYes, signed overflow can make f+1 > f be false. However in C/C++, because signed overflow is undefined behavior, the compiler is perfectly allowed to simplify the expression f+1 > f to always be true.
- stephencanon 7y agoThe compiler can make that simplification, but there's no guarantee that every compiler does, so you cannot assume that it will do so; it will sometimes happen to be true in a specific compiler, but it is never true in the standardized language model (which is what programs are written against).
- yoz-y 7y agoWhich would then have the same effect as the title's article, a bug with -O0 but working with -O3
- BubRoss 7y agoIt _could_ have that effect, but it would depend on signed integers wrapping.
- kbp 7y agoUnsigned integers are still specified to wrap.
- andrepd 7y agoi+1>i breaks even for unsigneds
- nybble41 7y agoAt least for unsigned integers i+1!=i, and the only exception to i+1>i is i==UINT_MAX. For signed integers it may be undefined behavior, of course—again in just one instance—but for floating point you may have f+1==f, or (f+1)-f>1 depending on the rounding mode, for many values of f which are nowhere near the limits of the range, not to mention other oddities like (f==f)==false when f is NaN. If you code involves floating-point operations then the number of corner cases you need to test increases greatly.
- MiroF 7y agoAlthough it's at least not UB
- patrec 7y agoWell, it may seem fairly obvious but it's also wrong.
- stephencanon 7y agoPerhaps it should be, but every 53 bit integer is exactly representable in double, because there's an implicit leading significand bit in the representation. It's also worth noting that every finite double with magnitude larger than 2^52 has a precise integer value; it's just that once you get beyond 2^53, not every integer is representable.
- ThreeFx 7y agoYes you're right - thanks for clarifying. I meant to say that not every integer in magnitude greater than 52 bit has an exact floating point representation in IEEE doubles.
- aidenn0 7y agoIIRC not every 53-bit integer is representable in double, since float has two zero representations, but twos-complement integers have only one. [edit] Since the extra value is precisely a power of 2 (-2^52), then it will round correctly, however the value is arguably not precisely -2^52 since it has an epsilon of greater than 1.
- stephencanon 7y agoThe spacing of floating point numbers have any bearing on their values. To be precise: every integer, positive or negative, with magnitude less than 2^53+1, is exactly represented in double-precision. “Extra values in two’s complement” don’t (and couldn’t possibly) effect this at all, since it is a statement about abstract integers and floating-point numbers, neither of which depends on two’s complement representations. In particular, -2^52 has a sign field of 1, an exponent field of 1023+52=1075, and an all-zero significand field. This number is exactly -2^52.
- qznc 7y agoSo the correct implementation of fits_long would be this? int fits_long(double value) { double max = (double) (1L << 52) double min = (double) -max; return min <= value && value <= max; } (assuming a 64bit long)
- syockit 7y agoWhen doing floating point arithmetic on the x86 though, it can extend to `80 - 1 (sign) - 15 (exponent) = 64` bits. So if the result of a floating point just so happens to have zero exponent, the mantissa can fit just right in a long int.
- stephencanon 7y agoThis hasn't been dependably true on x86 for almost two decades. SSE2 does double-precision computation at native width, not in the extended 80-bit format. Some 32b compilers still use 80-bit x87, but almost no 64b compilers do so.
- bitminer 7y ago> This hasn't been dependably true on x86 for almost two decades. SSE2 does double-precision computation at native width, not in the extended 80-bit format. Some 32b compilers still use 80-bit x87, but almost no 64b compilers do so. My experience has been different -- forcing SSE instructions gives me a different result on some math calculations. Core2 cpu, boost odeint calculations. Clang or gcc. Do you have a reference for why it's rare?
- stephencanon 7y agox87 is the default for 32b processes on Windows and Linux. 32b processes on macOS and 64b processes on all three OSes use SSE2 for double precision.