4 ms·
I'm not sure what you're trying to imply here. I'm not sure if it's relevant either, as this is about the implementation of the same instruction set (AVX/AVX2)
by tagrun 7y ago
I'm not sure what you're trying to imply here. I'm not sure if it's relevant either, as this is about the implementation of the same instruction set (AVX/AVX2) on Intel and AMD, whereas you say "not Intel <-> AMD though". Can you be more specific?
In any case, I never saw any real reason which warrant disabling an entire instruction set, such as AVX. Which is why you don't see such artificial crippling in open source implementations of LAPACK/BLAS/sundials/etc, and people (including me) have been using the same fortran code for many decades across many architectures.
And in case this is what you're trying to imply, no, they don't really give different numerical results on different CPUs.
- creato 7y agoMy relevant experience is in porting numerical code between CPUs and GPUs. Some of the issues that have caused problems are: - Different precision of approximate math (transcendental functions, reciprocals, etc.) - Different rounding behavior of the intermediate result in multiply-add instructions. - Different handling of exception cases (inf, nan, etc.). - Aside from correctness differences, some optimization strategies that make things faster on one processor make them slower on another. This happens even within different generations of x86 hardware. > Which is why you don't see such artificial crippling in open source implementations of LAPACK/BLAS/sundials/etc Are they as fast as MKL? If so, just use them? If not, why not? Maybe the reason is you can do better if you optimize for specific CPUs, with different latencies of various instructions?
- Dylan16807 7y agoPorting between different instruction sets is a very different thing from this situation. > Aside from correctness differences, some optimization strategies that make things faster on one processor make them slower on another. This happens even within different generations of x86 hardware. This is the one notably relevant part and, yeah, that's fine. Follow the CPUID features. Nobody expects it to be absolutely optimal on AMD. But let it use the code that was optimized for Intel chips with the same features.
- tagrun 7y agoThen it truly is irrelevant! You're not even talking about CPUs vs CPUs. The differences you are quoting coming from the difference in libraries (sin, exp etc will give different results depending on them libm implementation, that's normal and it has nothing to do with CPU instructions!), not the implementation of IEEE instructions (assuming that you're talking about IEEE floats, otherwise, you shouldn't expect them to behave the same in the first place!), though. > Are they as fast as MKL? If so, just use them? I (and a lot of other people) do use them, when I have a choice. Sometimes they are faster, sometimes they aren't. When there is a significant disparity, however, it usually is because of GeniueneIntel checks. > If not, why not? Because scientific software geared toward applications is usually closed-source proprietary or too complicated to be modified (remember that users aren't interested in becoming software engineers, in addition to their own jobs as researchers) to add new alternative backends and you don't get to choose.
- creato 7y ago> Then it truly is irrelevant! You're not even talking about CPUs vs CPUs. Why does that matter? The bulk of the issues come from implementation defined behavior, of which there is plenty within x86 itself to cause issues. In general, the IEEE-compliant parts of x86 are also IEEE-compliant on other processors, at least the ones I've dealt with. It's the operations that aren't specified by IEEE that cause problems.
- fock 7y ago"Sometimes they are faster, sometimes they aren't. When there is a significant disparity, however, it usually is because of GeniueneIntel checks." so you are saying that your open-source BLAS/LAPACK is showing performance differences (and worse performance compared to MKL) because of "something Intel". Seems like a lot of people here (including the ones not being able to compile numpy against another BLAS) are a little bit short on actual experience/knowing about the problem... "scientific software geared toward applications is usually closed-source proprietary or too complicated to be modified" If it's geared towards applications, it's usually opaque engineering stuff and the results of people claiming to do science with this software are mediocre at best... In my domain (qunatum-chemistry) nearly all software is delivered as source-distribution. Because modifications of methods are part of science...