11 ms·
An Interview with the Old Man of Floating-Point (1998)
- wheresmycraisin 5y agoHe taught numerical analysis at Berkeley, and though he was a great guy, I think he was waay to smart to be teaching undergrads... he'd go off on examples about literally every way that things like SVD could go wrong b/c of FP quirks, or how Matlab implements thing incorrectly, etc.
- kazinator 5y ago> I think it is nice to have at least one example -- IEEE 754 is one -- where sleaze did not triumph. A != A if A is a NaN: that's pretty sleazy. https://en.wikipedia.org/wiki/Law_of_identity https://en.wikipedia.org/wiki/Law_of_identity
- GeorgeTirebiter 5y agoWhat would you do? The comparison has no meaning if NaN.
- kazinator 5y agoIf it has no meaning, then the comparison should have an unspecified result, not true. Otherwise it has a meaning: the meaning of producing true! However, it is a poorly considered meaning which requires a thing to be different from itself. Since NaN values are valid representations which play a role in the system, and can be used in operations (such as comparing a NaN to a number, which is false), each of them must compare equal to itself. If the bits on the left are the same as the bits on the right, the comparison is true. Distinct NaN bit patterns are unequal. Simple as that. Whatever I would do, I would make sure that a comparison observes the Law of Identity. (I'd rather not have Inf and NaN at all; operations should just generate an exception if they can't come up with a number.)
- deleted 5y ago[deleted]
- wrs 5y agoDo you also think NULL should equal NULL in SQL? Read that Chesterton quote somebody posted above. It sounds like you need to study the standard a bit more, because you don’t seem to understand why NaNs are the way they are. You’re arguing from principles that don’t apply.
- kazinator 5y agoI don't know enough to say; all I remember from the few interactions I've had with SQL over my long programming career is that it's intellectually unsavory as a whole. I do think NIL should be EQ to NIL in Lisp.
- kevin_thibedeau 5y agoYou must like dividing by zero and never knowing about it. There's a reason why NaN blows things up. It's by design so that math errors don't propagate everywhere.
- kazinator 5y agoNo; I want an exception. The software image dies, unless it's handed.
- setr 5y agoYou have a fairly complex logic being stuffed into a binary operation — NaN == NaN and NaN != NaN are both irresponsible. The same with comparing to INF. The correct answer is that boolean operations don’t successfully represent the possibilities, and shouldn’t be offered in the first place. It’s the same with NULL in SQL.
- kazinator 5y agoThat is a valid view. NaN is supposed to propagate an error value, and that concept should continue through Boolean expressions. So that is to say, there has to be a NaT (not a truth) value which results instead of true or false, if a NaN is involved in a relational expression. Problem is, that is impractical. Programming languages tend to have two-valued Boolen baked into their DNA; it's implicit in if/then/else conditionals which will have to treat NaT as false --- back to square one. Programming languages with two-valued Booleans are not going to accommodate such a thing (it is not as easy to sneak in as NaN into floating-point). Even if they were to, programmers are going to be reluctant to turn every if/then situation into a three-way switch.
- B1FF_PSUVM 5y agoWhy? All we know it that both are not-numbers. It's just a label for something that cannot usefully be further identified. I'm pretty sure mathematicians can come up with many different not-numbers that have to share the same label, with a rather slim chance of being equal ...
- kazinator 5y agoIf X and Y are the same label, they should compare equal to satisfy the Law of Identity. If you perform two calculations, whose values are coming up as equal, but you didn't check the fact that both calculated the same NaN, that is your problem. Suppose you are looking for the result of the two calculations being unequal, and they produce NaN (or at least one of them does). That's also false positive. There is no way to get around checking for NaN.
- wrs 5y agoNumeric equality is not the same as bitwise equality. Use the right operation for your intention.
- hvdijk 5y agoNo one was talking about bitwise equality though. Bitwise equality means positive and negative zero do not compare as equal, and some NaNs compare as equal while others don't, in ways that will vary according to exactly where those NaNs came from. There may be times where that might be useful but it would be uncommon.
- kazinator 5y agoNothing which cheerfully concludes that two bitwise-identical operands are different can be called equality. The equality operation should only apply its own specific logic to a pair of operands which fail the bitwise test. (That doesn't mean all bits have to be looked at, like if an object has padding bits that don't contribute to the value.) In floating-point, like IEE754, a positive and negative zero still compare equal; they failed the bitwise test, so then the numeric logic still concludes they are the same number.
- jacobolus 5y ago> A != A if A is a NaN Among other things, this is super useful insofar as it gives a reliable way to test for NaN.
- kazinator 5y agoUsable and reliable is not the same thing as good design or good idea. The fact that Windows reserves the PRN file name in every directory is probably usable and reliable to someone. That doesn't mean it's a good design for providing access to a printer device. To test a set membership property of an object, you want a predicate function which takes the object as its only argument. For instance isnan(x). This predicate can be efficiently implemented without relying on a bastardized equality operation.
- kazinator 5y agoConsider that the compiler cannot optimize A != A into false, or A == A into true, because NaN values can occur at run time. While you might not explicitly write A == A into your code, it could occur implicitly due to some macro expansion, inline expansion or other code transformation. I think GCC with -ffast-math gets rid of this NaN rule and does such optimizations anyway. (Your code just has to avoid generating NaNs so that the optimizations are valid.)
- dang 5y agoA couple threads from way back: An Interview with the Old Man of Floating-Point (1998) - https://news.ycombinator.com/item?id=7769303 https://news.ycombinator.com/item?id=7769303 - May 2014 (17 comments) An Interview with the Old Man of Floating-Point (1998) - https://news.ycombinator.com/item?id=6656197 https://news.ycombinator.com/item?id=6656197 - Nov 2013 (21 comments)
- SavantIdiot 5y agoRandom anecdote: When I built my first 386 box that had a socket for a 387, I was super eager to fill that socket because even then PC builders were the same as today... but realized there wasn't any software that I used which would utilize it (my QuickC C-compiler didn't even support it!) The first app I remember that used it was Excel. It wasn't till the 486 that commodity games started using it.
- mistrial9 5y agoByte magazine !
- an1sotropy 5y agoWhen I teach about floating point, the two things I try to impress on the students are: it remains a truly incredible engineering feat to believably fit the entire real number line (plus infinities) into 32 or 64 bits, and, it was an incredible political feat to get so many competing companies to agree on one particular way of doing this; both are thanks to Kahan's leadership. Complaints about the quirks of using floating point could be tempered with some appreciation of the hard design decisions that were made, and with gratitude for the people who pulled it off.
- FabHK 5y agoYes. And one take-away for me: when yet another article comes out along the lines of "floating point sucks, and here's a much simpler and better replacement", and the author doesn't mention Kahan and shows in detail that they understand the design tradeoffs and decisions made back then (in IEEE 754), then there's a very good chance that you can toss it.
- dnautics 5y agoA lot of the design tradeoffs are not really relevant anymore[0]. There are some ways in which 754 effectively makes a "this is UB, up to the manufacturer" choice (to appease manufacturers of the day) which these days would probably not fly; it's a much easier sell to declare "no ub" (or the equivalent for hw) because we have retrospective power over all the times those were problems, and the hw manufacturers have far LESS power than the application consumer these days. [0] for example iirc cray had a wonky multiplier, don't remember if it was 754, that (I guess) they thought made it faster but resulted in noncommutative multiplication for many cases.
- FabHK 5y agoFavorite quotes: > members of the committee, for the most part, were about equally altruistic. IBM's Dr. Fred Ris was extremely supportive from the outset even though he knew that no IBM equipment in existence at the time had the slightest hope of conforming to the standard we were advocating. It was remarkable that so many hardware people there, knowing how difficult p754 would be, agreed that it should benefit the community at large. If it encouraged the production of floating-point software and eased the development of reliable software, it would help create a larger market for everyone's hardware. This degree of altruism was so astonishing that MATLAB's creator Dr. Cleve Moler used to advise foreign visitors not to miss the country's two most awesome spectacles: the Grand Canyon, and meetings of IEEE p754. > In the usual standards meetings everybody wants to grandfather in his own product. I think it is nice to have at least one example -- IEEE 754 is one -- where sleaze did not triumph. CDC, Cray and IBM could have weighed in and destroyed the whole thing had they so wished. Perhaps CDC and Cray thought `Microprocessors? Why should we worry?' In the end, all computer systems designers must try to make our things work well for the innumerable ( innumerate ?) programmers upon whom we all depend for the existence of a burgeoning market. > Epilog: The ACM's Turing award went to Kahan in 1989.
- bsder 5y agoThis is quite the whitewashing of something that was super contentious. IEEE 754 is okay as a technical standard. I've seen more denormals as a place to stash data than I ever have as a computational result, but, meh. However, the political reason wasn't that everybody were being magically altruistic. IBM and DEC, especially, were killing it and were in no way going to allow the other to set the standard. And everybody else was keen to stop IBM or DEC from having the de facto standard which would cement their dominance further. For example, if I remember correctly, Cray arithmetic was notorious for never actually being anything compliant (something about their multipliers).
- adrian_b 5y agoI believe that you are right, and we are very lucky for these historical circumstances, when both IBM and DEC have preferred to support the Intel floating-point number format, rather than accept the format of their major competitor. The Intel format, which is due to William Kahan, but also to Jerome Coonen and John Palmer and a few others with lesser contributions, was a huge improvement over the IBM and DEC floating-point formats. I have written my first programs when I was in high-school, for some IBM mainframes and DEC PDP-11 computers. Then I have used the Microsoft compilers for several languages, for Z80 / Intel 8080, which used FP formats similar to those of DEC. The IBM and DEC formats were really ugly and writing programs for them included many pitfalls. When I began to use an IBM PC, with its much more foolproof FP format, all problems disappeared. In general all processors introduced by Intel since their beginning and until the nineties (when their competition withered) lacked any innovative features. Every improvements in the early Intel CPUs had already been introduced earlier in CPUs from competitors. To this lack of innovation in CPU architecture, there is a very important exception, the IEEE 754 standard, which was based on the Intel format with very minor modifications. The Intel FP number format was one of the most important events in the history of floating-point numbers, the only other events with similar importance were when IBM made the first computers with hardware FP units (IBM 704 and NORC, in 1954) and when IBM introduced fused multiply-add (IBM POWER, in 1990).
- rodarmor 5y agoOne time me and a friend having an animated conversation on the 7th floor of Soda hall at Berkeley and William Kahan came out and gave us a coupon for Sizzlers. I think that was his way of telling us to get the fuck out.
- downut 5y agoAnd before our master Kahan, there was Pat H. Sterbenz. I still have my "Floating Point Computation" in the photocopied bound sheaf that was handed out to numerical analysis grad students at ASU in the late 80s. I learned an enormous amount about what digital computation means in the presence of algorithms in that class. EDIT: I have an 8087 chip always installed to sitting on my monitor base. Because Kahan.
- adgjlsfhk1 5y agoAs someone who has implemented a lot of low level functions using all the tricks of Floating point math, I have very mixed thoughts on Floating Point. Nan and -0.0 both seem like aggressively bad ideas to me. I can totally see why it was believed at the time that they would be good, but they just add a ton of special cases if you want to do things right that slow everything down. IMO, it would have been much better if we got errors instead of NaN (like we do for integer division by zero). That said, the ability to use double-double schemes to extend precision is wonderful and makes things much easier than they are in most of the Floating Point alternatives that have been proposed (eg Posits).
- an1sotropy 5y agoI'd heard of posits before but hadn't been motivated to learn about them until your comment; thanks. Reminds me of CIDR for IPv4
- adgjlsfhk1 5y agoIMO, 8 and 16 bit posits are way better than the float equivalents, but for 32 and 64 bits, floating point math is easier to analyze.
- adrian_b 5y agoAnyone who does not like NaNs should remember that using NaNs is just an option, nobody forces you to use them. It is enough to enable exception generation for undefined operations and it is guaranteed that no NaNs can appear in any results. It turns out that most people think that it is too much work to write suitable exception handlers, so they prefer to mask all exceptions and deal with NaNs and negative zeros. In practice this just means that you must be careful when you write conditional expressions where floating-point numbers are compared, because the order relation becomes partial, so there are 14 possible relational operations instead of the only 6 that are possible for a total order, and you also must write the compared values in a way where the sign of a zero does not matter (in most expressions the sign of a zero does not influence the result; you must make some efforts to distinguish a negative zero from a positive zero).
- aidenn0 5y agoNote that I have worked with chips designed in this century that did not implement denormals in hardware.
- jlgustafson 5y agoI hope I'm not too late to the party to correct some things I see here. The big accomplishment of Kahan and IEEE 754 was to get companies to agree on where the sign, exponent, and fraction should go, so that data interchange finally became possible between different computer brands. Kahan wanted decimal floats, not binary, and he wanted Extended Precision to be 128, not 80. I've had many hours of conversation with the man about how Intel railroaded that Standard to express the design decisions that had already been made for the i8087 coprocessor. John Palmer, who I also worked with for years, was proud of this, and told me "Whatever the i8087 is, THAT is the IEEE Standard." Posits have a single exception value, Not a Real (NaR) for all things that fall through the protections of C and Java and all the other modern languages for things like division by zero, and the square root of a negative value. Kahan wanted the quadrillions of Not a Number (NaN) patterns to be used to encode the address of the instruction in the program to pinpoint where it happened, but the support for this in computing languages never happened. By around 2005, vendors noticed they could trap the exceptions and spend hundreds of clock cycles handling them with microcode or software, so the FLOPS claims only applied to normal floats, not subnormals or NaN or infinities, etc. This is true today for all x86 and ARM processors, and SPARC for that matter. Only the POWER series from IBM can still claim to support IEEE 754 in hardware; hardware support for IEEE 754 is all but extinct. There are over a hundred papers published comparing posits and floats, both for accuracy on applications and difficulty of implementation. LLNL and Oxford U have definitively showed that posits are much more accurate than floats on a range of applications, so much so that a lower (power-of-two) precision can be used. Like 32-bit posits instead of 64-bit floats for shock hydrodynamics, and 16-bit posits instead of 32-bit floats for climate and weather prediction. For signal processing, 16-bit posits are about 10 dB more accurate (less noise) than 16-bit floats, which means they can perform lossless Fast Fourier Transforms (FFTs) on data from 12-bit A-to-D convertors. For the same precision, posit hardware add/subtract units appear slightly more expensive than float add/subtract, and multiplier units are slightly cheaper for posits than for floats. This echos what was found comparing the speed of the Berkeley SoftFloat emulator with that of Cerlane Leong's SoftPosit emulator. Naive studies say posits are more expensive because they first decode the posit into float subfields, apply time-honored float algorithms, then re-encode the subfields into posit format. This does not exploit the perfect mapping of posits to 2's complement integers. Float comparison hardware is quite complicated and expensive because there are redundant representation like –0 and +0 that have to test as equal, and redundant NaN exceptions that have to test as not equal even when their bit patterns are identical. Posit comparison hardware is unnecessary because they test exactly the same way as 2's complement integers. NaR is the 2's complement integer that has no absolute value and cannot be negated, 1000...000 in binary. It is equal to itself and less than any real-valued posit. The name is NaR because IEEE 754 incorrectly states that imaginary numbers are not numbers, and sqrt(–1) returns NaN. The Posit Standard is more careful to say that it is not a _real_. The Posit Standard is up to Version 4.13 and close to full approval by its Working Group. Don't use any Version 3 or earlier. The one on posithub.org may be out of date. In Version 4, the number of eS bits was fixed at 2, greatly simplifying conversions between different precisions. Unlike floats, posit precision can be changed simply by appending bits or rounding them off, without any need to decode the fraction and the scaling. It's like changing a 16-bit integer to a 32-bit integer; it costs next to nothing, which really helps people right-size the precision they're using.