6 ms·
I completely disagree. Your plot doesn't show CPU instructions are unpredictable (which is not correct). Yes, IEEE floats are not the same as real numbers (whic
by tagrun 7y ago
I completely disagree. Your plot doesn't show CPU instructions are unpredictable (which is not correct). Yes, IEEE floats are not the same as real numbers (which is what your plot is showing): they have inherent precision errors, and float operations are not even associative.
That being said, IEEE floats are carefully and consistently defined, and are perfect predictable. The unpredictability you claim is not due to stochastic errors or faulty implementation of CPU vendors, they're a part of the IEEE definition and are deterministic.
- creato 7y agoThe '1/x' part of this experiment is described by an IEEE spec. The reciprocal instructions are not. This experiment doesn't cover the difference between theoretical real numbers and float computations, it's the difference between two different ways of working with floats. One of them is IEEE specified, the other is not. In practice, using these kinds of instructions (which are not specified by IEEE) can give massive performance advantages.
- tagrun 7y agoAnd how does the difference between IEEE-754 1/x and SSE reciprocal instruction on the same CPU is relevant to Intel disabling AVX/AVX2 on AMD CPUs? So your logic is: given that IEEE-754 1/x and SSE rcpps yield different results on a single Intel CPU, then... AMD's AVX implementation cannot be the same as Intel's AVX implementation and therefore, Intel is perfectly correct in disabling AVX on AMD?
- creato 7y agoThe logic is: 1. Some instructions definitely have implementation defined behavior that will vary across platforms. rcpps is just one example to illustrate the point. 2. Therefore, if you test something on an intel CPU, it might not behave the same on an AMD CPU (and vice versa). This doesn't mean AMD's implementation is bad. (For all you know, it might work on intel because you are accidentally exploiting a bug, and that bug won't exist on AMD, and break your software.) 3. Therefore, if you don't test something using these on an AMD CPU, it could be risky to enable it. As an aside, I don't necessarily think this logic is the only factor that should go into this decision and I'm not sure it's the right one for intel to make. But there is some logic to it that isn't simply intel screwing over AMD, which is the only point I've been trying to make here. I think what you can fairly say about intel here is they are being lazy and overly risk averse, but not anti-competitive (in this instance!).
- greglindahl 7y agoThis is a pretty long thread for you to not notice that your data doesn't show an implementation defined behavior difference because it's ... a graph of just one implementation. 1/x has been around as an approximation used in the overwhelming majority of floating point division units for a long time, even before the SSE instruction came about, so I'm pretty dubious that either Intel or AMD is changing their answers. Perhaps you could prove it.
- Const-me 7y agoHe’s right. The values are actually different. Here’s a graph of two: http://const.me/tmp/vrcpps-errors-chart.png http://const.me/tmp/vrcpps-errors-chart.png AMD is Ryzen 5 3600, Intel is Core i3 6157U. Over the complete range of floats, AMD is more precise on average, 0.000078 versus 0.000095 relative error. However, Intel has 0.000300 maximum relative error, AMD 0.000315. Good news is both are well within the spec. The documentation says “maximum relative error for this approximation is less than 1.5*2^-12”, in human language that would be 3.6621E-4. Source code: https://gist.github.com/Const-me/a6d36f70a3a77de00c61cf4f6c17c7ac https://gist.github.com/Const-me/a6d36f70a3a77de00c61cf4f6c1...
- magicalhippo 7y agoThe LHC@Home BOINC project had issues due to this. Work units calculated on Intel CPUs would frequently not match the same work unit calculated on AMD CPUs. They traced it down to the subtle difference in results from certain instructions, like shown in parent. Since the LHC@Home project simulated protons circulating in the LHC, the small differences added up over the typically large number of time steps. IIRC their solution was to pay a small price and use software implementations of those instructions, but I can't find the reference right now.
- tagrun 7y agoThat indicates a numerically unstable code on his/her part though. Scientific results shouldn't depend on unspecified epsilon values that fall within the spec. If they're getting different results on different CPUs which both implement the spec correctly, they need to fix their code anyways.