3 ms·
And how does the difference between IEEE-754 1/x and SSE reciprocal instruction on the same CPU is relevant to Intel disabling AVX/AVX2 on AMD CPUs? So your lo
by tagrun 7y ago
And how does the difference between IEEE-754 1/x and SSE reciprocal instruction on the same CPU is relevant to Intel disabling AVX/AVX2 on AMD CPUs?
So your logic is: given that IEEE-754 1/x and SSE rcpps yield different results on a single Intel CPU, then... AMD's AVX implementation cannot be the same as Intel's AVX implementation and therefore, Intel is perfectly correct in disabling AVX on AMD?
- creato 7y agoThe logic is: 1. Some instructions definitely have implementation defined behavior that will vary across platforms. rcpps is just one example to illustrate the point. 2. Therefore, if you test something on an intel CPU, it might not behave the same on an AMD CPU (and vice versa). This doesn't mean AMD's implementation is bad. (For all you know, it might work on intel because you are accidentally exploiting a bug, and that bug won't exist on AMD, and break your software.) 3. Therefore, if you don't test something using these on an AMD CPU, it could be risky to enable it. As an aside, I don't necessarily think this logic is the only factor that should go into this decision and I'm not sure it's the right one for intel to make. But there is some logic to it that isn't simply intel screwing over AMD, which is the only point I've been trying to make here. I think what you can fairly say about intel here is they are being lazy and overly risk averse, but not anti-competitive (in this instance!).
- greglindahl 7y agoThis is a pretty long thread for you to not notice that your data doesn't show an implementation defined behavior difference because it's ... a graph of just one implementation. 1/x has been around as an approximation used in the overwhelming majority of floating point division units for a long time, even before the SSE instruction came about, so I'm pretty dubious that either Intel or AMD is changing their answers. Perhaps you could prove it.
- Const-me 7y agoHe’s right. The values are actually different. Here’s a graph of two: http://const.me/tmp/vrcpps-errors-chart.png http://const.me/tmp/vrcpps-errors-chart.png AMD is Ryzen 5 3600, Intel is Core i3 6157U. Over the complete range of floats, AMD is more precise on average, 0.000078 versus 0.000095 relative error. However, Intel has 0.000300 maximum relative error, AMD 0.000315. Good news is both are well within the spec. The documentation says “maximum relative error for this approximation is less than 1.5*2^-12”, in human language that would be 3.6621E-4. Source code: https://gist.github.com/Const-me/a6d36f70a3a77de00c61cf4f6c17c7ac https://gist.github.com/Const-me/a6d36f70a3a77de00c61cf4f6c1...
- magicalhippo 7y agoThe LHC@Home BOINC project had issues due to this. Work units calculated on Intel CPUs would frequently not match the same work unit calculated on AMD CPUs. They traced it down to the subtle difference in results from certain instructions, like shown in parent. Since the LHC@Home project simulated protons circulating in the LHC, the small differences added up over the typically large number of time steps. IIRC their solution was to pay a small price and use software implementations of those instructions, but I can't find the reference right now.
- tagrun 7y agoThat indicates a numerically unstable code on his/her part though. Scientific results shouldn't depend on unspecified epsilon values that fall within the spec. If they're getting different results on different CPUs which both implement the spec correctly, they need to fix their code anyways.
- magicalhippo 7y agoBesides using a deterministic floating-point implementation, how can you avoid that in sims like this? A slightly difference force on the particle at current time step will cause it to end up in a slightly different place. edit: They do run sims with slightly different initial conditions to get the physics rather than simulation artifacts. The issue here is that for a given set of initial conditions, results computed on different machines would not agree.
- tagrun 7y agoThis makes no sense at all. The differences that you're seeing fall within the spec. You're basically saying that Intel is justified in crippling AMD CPUs if AMD's floating point instructions doesn't implement Intel's out-of-spec quirks as well. This is a flimsy argument, the FUD that Intel has been propagating for years now, and I'm not sure why you keep on pushing the Intel FUD so hard again and again. There is basically nothing which guarantees that the numerical errors that are within the error tolerance of the spec won't change in new iterations of Intel CPUs either. If the correctness of your calculation depends on the values of the unspecified within the spec, this means you need to change the code to a more stable algorithm or use higher precision floats anyway.
- creato 7y ago> There is basically nothing which guarantees that the numerical errors that are within the error tolerance of the spec won't change in new iterations of Intel CPUs either. You're absolutely right. The difference is that if a customer of a library like MKL comes to intel and reports a problem like this, intel is going to help troubleshoot and fix the problem, and they're going to have a lot of internal technical documentation to help understand and fix the differences. > The differences that you're seeing fall within the spec. ... This is a flimsy argument, the FUD that Intel has been propagating for years now, and I'm not sure why you keep on pushing the Intel FUD so hard again and again. The problem is that it's really hard to know if code depends only on the requirement in the spec, or depends on more specific behavior present in the machine(s) you tested on. Call it FUD if you want. A few people in this thread have reported spending significant effort on troubleshooting implementation defined behavior. Yeah, code that depends on this is bad, but bad code exists and sometimes people with the bags of money care about it.