4 ms·
Out of curiosity, I started watching this talk, but right there in the introduction, the very first example is a dot product where supposedly IEEE 754 double pr
by sunfish 10y ago
Out of curiosity, I started watching this talk, but right there in the introduction, the very first example is a dot product where supposedly IEEE 754 double precision gets the wrong answer; I stopped to check the result, and I got the correct answer with double precision (even without binary sum collapse). Then, he says the x87's results are nondeterministic due to being affected by cache, which is incorrect. Then, he trots out a series of half-truths about IEEE 754, which are somewhat true but don't mean what he's presenting them to mean. He even cites a Cray 1 quirk as an example of IEEE 754 weirdness, when in reality, the Cray 1 (1975) predated IEEE 754 by about 10 years.
Designing a good arithmetic system takes a lot of attention to detail, and he doesn't make a good impression by being that loose with details in his introduction.
- dbcurtis 10y agoThe Cray 1 was done before anyone realized the importance of denormals. And denormals are gruesome to implement on vector machines anyway.
- greglindahl 10y agoOnly a subset of scientific codes benefit from denormals, and it does not appear that implementing them is that big of a deal, given that 4-stage pipelines can do it. The pipeline depth to memory is a lot bigger than 4.
- dbcurtis 10y agoIf you consider every iterative algorithm that solves systems of nonlinear equations a subset you cam ignore.... SPICE, linpack, fishpack.... sorry, denormals are essential to today's algorithms. As to "no big deal".... show me your code. Don't worry, I'll be able to understand it. After 20 years as a CPU designer I've learned how to understand bit-bashing.
- greglindahl 10y agoThe iterative algorithms for all of the scientific fields that I've worked on don't fall within the subset that you're describing, so no, I'm not going to agree that you're speaking about the majority of today's algorithms. Weather, climate, finite differencing, FFT... nope. Implicit systems, sure. Radar cross-sections, sure. And as a CPU designer, are you disagreeing that most FP units today take 4 cycles (fully pipelined) (add and multiply, of course)? And that main memory is a lot farther away than that? Given 20 years of experience, you missed out on the Cray PVP machines that were the start of this sub-thread. But the cycle counts I'm giving are modern ones. Cray did eventually implement IEEE in these machines without any significant problems, but that was more than 20 years ago.
- nkurz 10y agoOnly a subset of scientific codes benefit from denormals, and it does not appear that implementing them is that big of a deal, given that 4-stage pipelines can do it. Maybe I'm misunderstanding your terminology, but it seems like you are saying that operations involving denormals have the same latency as normal floating point multiplications and additions. At least for multiplication on Intel chips through Haswell, I think the current case is that subnormals between 0 and FLT_MIN still have abysmal performance --- 100+ cycles of penalty. Here's Bruce Dawson from a few years ago: Performance implications on SSE Intel handles NaNs and infinities much better on their SSE FPUs than on their x87 FPUs. NaNs and infinities have long been handled at full speed on this floating-point unit. However denormals are still a problem. On Core 2 processors the worst-case I have measured is a 175 times slowdown, on SSE addition and multiplication. On SandyBridge Intel has fixed this for addition – I was unable to produce any slowdown on ‘addps’ instructions. However SSE multiplication (‘mulps’) on Sandybridge has about a 140 cycle penalty if one of the inputs or results is a denormal. https://randomascii.wordpress.com/2012/05/20/thats-not-normalthe-performance-of-odd-floats/ https://randomascii.wordpress.com/2012/05/20/thats-not-norma... And here's the overview from a slightly more recent paper: C. Subnormal Performance Variability Due to the complex nature of the floating point numbers, processors struggle to handle certain inputs efficiently. In particular, it is well understood that operating on subnormal values can cause extreme performance issues, including slowdowns of up to 100× [19]. As an example, on a Core i7 processor using SSE instructions, performing standard mul- tiply between two normal numbers takes 4 clock cycles, whereas the same multiply given a subnormal input takes over 200 clock cycles. https://cseweb.ucsd.edu/~hovav/dist/subnormal.pdf https://cseweb.ucsd.edu/~hovav/dist/subnormal.pdf Are you saying this has been fixed in recent (or non-Intel) chips? Or maybe you were considering only NaN and Inf when you said denormals?
- greglindahl 10y agoSorry, you're correct, I was thinking about the speed of NaN and Inf.
- effie 10y agoDoubles on my Intel give correct answer too, but change the exponent used everywhere to 8 and you'll see the problem. Gustafson may have used different implementation of IEEE 754 that gave him different result. There are problems with repeatability of float operations such as loss of precision when moving value from registers to memory and data alignment issues, but it seems that these can be avoided with proper compiler options, at the expense of speed. https://gcc.gnu.org/bugzilla/show_bug.cgi?id=323 https://gcc.gnu.org/bugzilla/show_bug.cgi?id=323 https://software.intel.com/en-us/articles/run-to-run-reproducibility-of-floating-point-calculations-for-applications-on-intel-xeon https://software.intel.com/en-us/articles/run-to-run-reprodu...
- sunfish 10y agoThere is no "different implementation of IEEE 754" that could give a different result in double precision for that example. There are reasonable criticisms of IEEE 754; that's not the issue.
- effie 10y ago> There is no "different implementation of IEEE 754" that could give a different result in double precision for that example. How do you know that?
- sunfish 10y agoIEEE 754 is a standard, and it requires a specific, bitwise-reproducible, answer for the computation in question. I'm the author of most of the floating-point tests in WebAssembly's conformance testsuite, which tests such things in practice across several hardware platforms. One of the half-truths in the presentation (in the intro) is that IEEE 754 is a mixture of requirements and recommendations. IEEE 754 does have both requirements and recommendations, however what the presentation doesn't say is that, within a given format like double precision (aka binary64), the basic operations like add, subtract, multiply, divide, squareRoot, etc.) have exactly one possible result for any given input (except that NaNs may have some implementation-defined bits, though this is usually unimportant).