6 ms·
I can't tell if you're being naive or deliberately perpetuating an Intel FUD. You're being part of the problem by doubling down on some ambiguous comment (which
by tagrun 7y ago
I can't tell if you're being naive or deliberately perpetuating an Intel FUD. You're being part of the problem by doubling down on some ambiguous comment (which is not even based on any real test) as if there actually is a problem with AMD's AVX implementation.
As a physicists, I can tell you for sure that no, this is not why Intel disables it. Intel had quite a success by purposely crippling icc/ifort and MKL on AMD, and created GenuineIntel as a legal barrier to prevent AMD (and AMD users) to come up with a workaround, which allowed Intel to essentially kill the competition in HPC.
It is a well known fact that Intel has been actively working to cripple AMD's performance across the board. They invented GenuineIntel exclusively for that purpose, which was a part of the bigger picture filled with false advertisements, bribes (the most famous Intel bribe cases involved Dell), smear campaigns, lawsuits, so on and so forth. Intel has a long history of playing dirty against competitors.
- creato 7y ago> I can't tell if you're being naive or deliberately perpetuating an Intel FUD. You're being part of the problem by doubling down on some ambiguous comment (which is not even based on any real test) I have never actually seen Intel comment on this, I just personally have experienced headaches porting numerical code between two different architectures (not AMD <-> Intel though), where on the surface the architectures appear to be the same and straightforward to swap between, but in practice they are not. This doesn't mean one of the architectures had "problems" and the other did not. It's not about one architecture being inferior than the other, they're simply different.
- deleted 7y ago[deleted]
- tagrun 7y agoI'm not sure what you're trying to imply here. I'm not sure if it's relevant either, as this is about the implementation of the same instruction set (AVX/AVX2) on Intel and AMD, whereas you say "not Intel <-> AMD though". Can you be more specific? In any case, I never saw any real reason which warrant disabling an entire instruction set, such as AVX. Which is why you don't see such artificial crippling in open source implementations of LAPACK/BLAS/sundials/etc, and people (including me) have been using the same fortran code for many decades across many architectures. And in case this is what you're trying to imply, no, they don't really give different numerical results on different CPUs.
- creato 7y agoMy relevant experience is in porting numerical code between CPUs and GPUs. Some of the issues that have caused problems are: - Different precision of approximate math (transcendental functions, reciprocals, etc.) - Different rounding behavior of the intermediate result in multiply-add instructions. - Different handling of exception cases (inf, nan, etc.). - Aside from correctness differences, some optimization strategies that make things faster on one processor make them slower on another. This happens even within different generations of x86 hardware. > Which is why you don't see such artificial crippling in open source implementations of LAPACK/BLAS/sundials/etc Are they as fast as MKL? If so, just use them? If not, why not? Maybe the reason is you can do better if you optimize for specific CPUs, with different latencies of various instructions?
- Dylan16807 7y agoPorting between different instruction sets is a very different thing from this situation. > Aside from correctness differences, some optimization strategies that make things faster on one processor make them slower on another. This happens even within different generations of x86 hardware. This is the one notably relevant part and, yeah, that's fine. Follow the CPUID features. Nobody expects it to be absolutely optimal on AMD. But let it use the code that was optimized for Intel chips with the same features.
- tagrun 7y agoThen it truly is irrelevant! You're not even talking about CPUs vs CPUs. The differences you are quoting coming from the difference in libraries (sin, exp etc will give different results depending on them libm implementation, that's normal and it has nothing to do with CPU instructions!), not the implementation of IEEE instructions (assuming that you're talking about IEEE floats, otherwise, you shouldn't expect them to behave the same in the first place!), though. > Are they as fast as MKL? If so, just use them? I (and a lot of other people) do use them, when I have a choice. Sometimes they are faster, sometimes they aren't. When there is a significant disparity, however, it usually is because of GeniueneIntel checks. > If not, why not? Because scientific software geared toward applications is usually closed-source proprietary or too complicated to be modified (remember that users aren't interested in becoming software engineers, in addition to their own jobs as researchers) to add new alternative backends and you don't get to choose.
- im3w1l 7y agoHow solid is that legal barrier? User agent strings which is a similar case are really weird to deal with sniffing. Edge supposedly has the string "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/78.0.3904.108 Safari/537.36 Edg/44.18362.449.0".