5 ms·
And how exactly does that plot show that AMD's AVX instructions are faulty?
by tagrun 7y ago
And how exactly does that plot show that AMD's AVX instructions are faulty?
- creato 7y agoIt doesn't. It shows that the behavior of instructions like this is very unpredictable and subtle, and that code using it that was tested on one processor might not behave the same on another processor. It doesn't mean one is faulty and the other is not.
- jchw 7y agoA processor would be faulty if the error or behavior fell outside of the architecture specification. The behavior is not unpredictable because the behavior is defined by a specification.
- tagrun 7y agoI completely disagree. Your plot doesn't show CPU instructions are unpredictable (which is not correct). Yes, IEEE floats are not the same as real numbers (which is what your plot is showing): they have inherent precision errors, and float operations are not even associative. That being said, IEEE floats are carefully and consistently defined, and are perfect predictable. The unpredictability you claim is not due to stochastic errors or faulty implementation of CPU vendors, they're a part of the IEEE definition and are deterministic.
- creato 7y agoThe '1/x' part of this experiment is described by an IEEE spec. The reciprocal instructions are not. This experiment doesn't cover the difference between theoretical real numbers and float computations, it's the difference between two different ways of working with floats. One of them is IEEE specified, the other is not. In practice, using these kinds of instructions (which are not specified by IEEE) can give massive performance advantages.
- tagrun 7y agoAnd how does the difference between IEEE-754 1/x and SSE reciprocal instruction on the same CPU is relevant to Intel disabling AVX/AVX2 on AMD CPUs? So your logic is: given that IEEE-754 1/x and SSE rcpps yield different results on a single Intel CPU, then... AMD's AVX implementation cannot be the same as Intel's AVX implementation and therefore, Intel is perfectly correct in disabling AVX on AMD?
- creato 7y agoThe logic is: 1. Some instructions definitely have implementation defined behavior that will vary across platforms. rcpps is just one example to illustrate the point. 2. Therefore, if you test something on an intel CPU, it might not behave the same on an AMD CPU (and vice versa). This doesn't mean AMD's implementation is bad. (For all you know, it might work on intel because you are accidentally exploiting a bug, and that bug won't exist on AMD, and break your software.) 3. Therefore, if you don't test something using these on an AMD CPU, it could be risky to enable it. As an aside, I don't necessarily think this logic is the only factor that should go into this decision and I'm not sure it's the right one for intel to make. But there is some logic to it that isn't simply intel screwing over AMD, which is the only point I've been trying to make here. I think what you can fairly say about intel here is they are being lazy and overly risk averse, but not anti-competitive (in this instance!).
- greglindahl 7y agoThis is a pretty long thread for you to not notice that your data doesn't show an implementation defined behavior difference because it's ... a graph of just one implementation. 1/x has been around as an approximation used in the overwhelming majority of floating point division units for a long time, even before the SSE instruction came about, so I'm pretty dubious that either Intel or AMD is changing their answers. Perhaps you could prove it.
- Const-me 7y agoHe’s right. The values are actually different. Here’s a graph of two: http://const.me/tmp/vrcpps-errors-chart.png http://const.me/tmp/vrcpps-errors-chart.png AMD is Ryzen 5 3600, Intel is Core i3 6157U. Over the complete range of floats, AMD is more precise on average, 0.000078 versus 0.000095 relative error. However, Intel has 0.000300 maximum relative error, AMD 0.000315. Good news is both are well within the spec. The documentation says “maximum relative error for this approximation is less than 1.5*2^-12”, in human language that would be 3.6621E-4. Source code: https://gist.github.com/Const-me/a6d36f70a3a77de00c61cf4f6c17c7ac https://gist.github.com/Const-me/a6d36f70a3a77de00c61cf4f6c1...