12 ms·
Intel's “Cripple AMD” Function (2019)
- DSingularity 4y ago> After Intel had flatly denied to change their CPU dispatcher, I decided that the most efficient way to make them change their minds was to create publicity about the problem. I contacted several IT magazines, but nobody wanted to write about it. Sad, but not very surprising, considering that they all depend on advertising money from Intel. Sorry to go on this tangent: but is capitalism so rotten that everything eventually corrupts? Here even outlets for discussion on topics of science and technology self-censure to maximize profit. So much for freedom of speech. Where is truth these days?
- BlueTemplar 4y agoStrictly speaking, it's not about "capitalism", but journalists trying to get funding from anywhere else than directly their readership. (There used to be law proposals forbidding this, not sure if advertising was already on their radar back in the 1940's.) Of course there would still remain the issue of self-censoring to avoid annoying your readership, not sure how you can deal with that...
- holdenk 4y agoHuh I had wondered why I saw so many Python packages blacklist MKL now I know why.
- dbcurtis 4y agoThe philosophy behind MKL is that each CPU vendor provides an MKL for their CPU. If you expect to mix and match MKLs and CPUs, you don’t understand the goals of MKL.
- monocasa 4y agoAre there any implementations of MKL other than Intel's?
- ReleaseCandidat 4y agoNo. There are AMD's AOCL and Apple's 'Accelerate', but of subsets of the MKL only AFAIK. https://developer.amd.com/amd-aocl/ https://developer.amd.com/amd-aocl/ https://developer.apple.com/documentation/accelerate https://developer.apple.com/documentation/accelerate
- stephencanon 4y agoAccelerate and MKL have some overlap (notably BLAS, LAPACK, signal processing libraries and basic vectorized math operations), but each also contains a whole bunch of API that the other lacks. Neither is a subset of the other. They both contain a sparse matrix library, but exactly what operations are offered is somewhat different between the two. They both have image processing operations, but fairly different ones. Accelerate has BNNS, MKL has its own set of deep learning interfaces...
- kxyvr 4y agoIn case you or anyone else knows, are there other libraries that implement a high performance sparse QR? Really, I need a Q-less QR factorization for sparse matrices. As far as I know, there are only two: one comes from MKL: https://www.intel.com/content/www/us/en/developer/articles/technical/intel-mkl-sparse-qr-solver-multifrontal-sparse-qr-factorization-method-for-solving-a-sparse.html https://www.intel.com/content/www/us/en/developer/articles/t... The other comes from SPQR, which is part of SuiteSparse: https://people.engr.tamu.edu/davis/suitesparse.html https://people.engr.tamu.edu/davis/suitesparse.html Part of the issue is that SPQR is dual licensed GPL/Commercial and the last time I checked a license was not cheap. Conversely, MKL has no redistribution fee, so it's been essentially the only option for this factorization if the code can't be bundled in a way compatible with the GPL.
- stephencanon 4y agoReplying to [dead] sibling post from kxyvr: yes, Accelerate provides a Q-less sparse QR on Apple platforms (https://developer.apple.com/documentation/accelerate/sparse_solvers https://developer.apple.com/documentation/accelerate/sparse_..., in particular SparseFactorizationCholeskyAtA). I believe that MA49 from HSL does it as well, and may have more acceptable licensing than SuiteSparse depending on your situation.
- wyldfire 4y agoEach CPU vendor or each CPU architecture? (genuinely asking, I don't know how it's intended)
- bee_rider 4y agoThe expectation in the HPC community is that an interested vendor will provide their own BLAS/LAPACK implementation (MKL is a BLAS/LAPACK implementation, along with a bunch of other stuff), which is well-tuned for their hardware. These sort of libraries aren't just tuned for an architecture, they might be tuned for a given generation or even particular SKUs.
- hallway_monitor 4y agoI learned about this recently when trying to optimize ML test architecture running on Azure. It turns out having access to Ice Lake chips would allow optimizations that should decrease compute time and therefore cost by 20-30%.
- bee_rider 4y agoSome AVX-512 stuff I guess? AVX-512 had a rough rollout, but it seems like it is finally turning into something nice.
- wmf 4y agoEach vendor. Intel BLAS (MKL) has Intel-specific optimizations and AMD BLAS has AMD-specific optimizations. Intel is still acting in bad faith by allowing MKL to run in crippled mode on AMD. They should either let it use all available instructions or make it refuse to run.
- microtonal 4y agoThe latest oneMKL versions have sgemm/dgemm kernels for Zen CPUs that are almost as fast as the AVX2 kernels (that require disabling Intel CPU detection on Zen).
- ReleaseCandidat 4y agoThat would be 'each CPU vendor provides an optimized BLAS library for their CPU'. The problem is that Intel's MKL is more than just BLAS. But AMD does have its own optimized libraries: https://developer.amd.com/amd-aocl/ https://developer.amd.com/amd-aocl/
- Kon-Peki 4y agoThis has been discussed on HN before. I don't condone Intel behavior, but let's be honest here: AMD underinvests in software and expects others to pick up the slack. That isn't acceptable.
- amelius 4y agoI think it's great if a hardware company leaves the software for others. This leads to open specifications.
- ethbr0 4y agoAt the firmware / driver level, fully open specifications for high performance hardware is an impossible dream. At best, detailed documentation is a lower priority item below "make it work" and "increase performance". At worst, it requires exposing trade secrets. Edit: It'd probably be more productive for everyone if we set incentives and work such that the goal we want (compilers that produce code that runs optimally on Intel, AMD, and other architectures) isn't contingent on Intel writing them for non-Intel architectures. (Said somewhat curmudgeonly, because everyone complains about things like this, but also doesn't really how insanely hard and frustratingly edge-case-ridden compiler work is)
- sedatk 4y agoNo, just don't falsely market your product as fair or neutral.
- dodobirdlord 4y agoIt’s the Intel MKL, I don’t think Intel has ever even endorsed using it on other vendors CPUs, much less claimed that it is “fair” or “neutral”.
- ReleaseCandidat 4y agoWell: On November 12, 2009 AMD and Intel Corporation announced a comprehensive settlement agreement to end all outstanding legal disputes between the companies, including antitrust and patent cross license disputes. In addition to a payment of $1.25B that Intel made to AMD, Intel agreed to abide by an important set of ground rules that continue in effect until November 11, 2019. Customers and Partners With respect to customers and partners, Intel must not:* [...] Intentionally include design/engineering elements in its products that artificially impair the performance of any AMD microprocessor. https://www.amd.com/en/corporate/antitrust-ruling https://www.amd.com/en/corporate/antitrust-ruling I like that 'in effect until November 11, 2019.' part :D
- mechanical_bear 4y agoSo it appears not only is this posting from 2019, but the most recent information they reference is 2010. This seems to be no longer relevant? I’d love it if submissions on HN had a small blurb from the author explaining why their submission is interesting/relevant.
- shrx 4y agoIt certainly is still relevant. Please read the article (and the 2020 update below) before commenting on it.
- Anunayj 4y agoyup, infact recently MATLAB applied a fix to this for their software [1] [1] https://www.extremetech.com/computing/308501-crippled-no-longer-matlab-2020a-runs-amd-cpus-at-full-speed https://www.extremetech.com/computing/308501-crippled-no-lon...
- MikePlacid 4y agoThank you for the link. It helped me to find the actual performance difference. It is significant: AMD’s performance improves by 1.32x – 1.37x overall… changing what looked like a narrow victory [for Intel] over the 3960X and a good showing against the 3970X into an all-out loss. https://www.extremetech.com/computing/302650-how-to-bypass-matlab-cripple-amd-ryzen-threadripper-cpus https://www.extremetech.com/computing/302650-how-to-bypass-m...
- mechanical_bear 4y agoI didn’t state it was not relevant, I asked. Glad to see there is an update. My main point stands though, I’d love to see an explanation with posts.
- nonplus 4y agoI would love to see a 2022 follow-up from Agner Fog on this. He has work on C++ compilers as recently as 2021 so I'm sure he has recent real world info on the topic. https://www.agner.org/optimize/#manuals https://www.agner.org/optimize/#manuals
- Certified 4y agoLast time this came up on Hacker News I discovered SolidWorks 2021 was using an older MKL library that supports the MKL_DEBUG_CPU_TYPE=5 environment variable. I'm on an AMD cpu and measured a small solidworks fps and rebuild time improvement with the flag enabled
- eatonphil 4y agoThe first comment seems to suggest that flag no longer works. https://www.agner.org/forum/viewtopic.php?t=6#p82 https://www.agner.org/forum/viewtopic.php?t=6#p82
- Certified 4y agoMultiple versions of MKL dlls exist in the install directory of Solidworks 2021. Indeed, the dlls supporting FloXpress and simulation seem to be the updated MKL version that no longer support the flag. However, the main executable only seems to call sldmkl_parts.dll. It appears to be MKL version 2018.1.156 that does support the flag
- bee_rider 4y agoIt would depend on the version of MKL. If Solidworks has (just for example) statically linked to or bundled in an old version of MKL, then it should work there, still.
- mhh__ 4y agoThis could arguably dated 2009 as that is when it was originally discovered (approx.). https://www.agner.org/optimize/blog/read.php?i=49 https://www.agner.org/optimize/blog/read.php?i=49
- ReleaseCandidat 4y agoIt has been discovered before that, at least in 2005: https://techreport.com/news/8547/does-intels-compiler-cripple-amd-performance/ https://techreport.com/news/8547/does-intels-compiler-crippl...
- midjji 4y agoSo blacklist intel compiler in favor or GCC and CLANG, seems entirely reasonable!
- nomel 4y agoThis assumes there's no performance/$$$ loss to switching, that you would have to justify to your org.
- harry8 4y agoFor anyone shipping binaries to customers using the Intel compiler could well be considered negligent. Intel have made it clear they will secretly sabotage /your/ customers if you use their tools to make your product and in fact they have done so. They will secretly sabotage you if you aren't a "pure intel" shop. Those actions were and remain completely hostile. In the light of "Reflections on Trusting Trust" [1] "Intel cannot be trusted to supply your compiler at any price." That's a point of view that is a lot more than just "a reasonable one to hold." The reflection on Intel and their lack of reckoning having been caught out sabotaging your customers is something any customer of Intel needs to consider - included in that assessment must be the expected value of the $$$ loss of purchasing from Intel. It's really not something anyone can responsibly ignore and fail to assess. Then go ahead and making your responsible and informed engineering and business trade off. [1] https://www.cs.cmu.edu/~rdriley/487/papers/Thompson_1984_ReflectionsonTrustingTrust.pdf https://www.cs.cmu.edu/~rdriley/487/papers/Thompson_1984_Ref... edit: The point being we all have a bar of "well they wouldn't actually do that" in a purchasing decision. That bar for Intel is dramatically lower as a result of this incident and failure to properly address it in full with a mea culpa and consequences rather than the ongoing minimum action required by the courts and damage limitation we've seen. It is very hard to see how the probability of them secretly sabotaging your goals could have gone down here. That is what Intel think of their reputation and what they think they can get away with in their response to you.
- midjji 4y agoI almost wish my org knew enough to make such policies.
- phkahler 4y agoWorth noting that Intel has dropped their "old" compiler and the newer "Intel" compilers are LLVM based. IMHO they will likely be pulling similar anti-AMD tricks with it and they are keeping their paid version closed source - which is allowed by LLVMs license. RMS was right that compilers should be GPL licensed to prevent exactly this kind of thing (and worse things which are haven't happened yet). On another compiler related note, I find it insane that GCC had not turned on vectorization at optimization -O2 for the x68-64 targets. The baseline for that arch has SSE2, so vectorization has always made sense there. The upcoming GCC 12 will have it enabled at -O2. I'd bet the Intel compiler always did vectorization at -O2 for their 64bit builds.
- mistrial9 4y ago> x68-64 targets thats a typo .. are you showing a case of AVX instructions not generated by GCC? where are the details here? Is SSE2 from twenty years ago? https://en.wikipedia.org/wiki/SSE2 https://en.wikipedia.org/wiki/SSE2
- phkahler 4y agoThe original AMD64 extension and associated ABI included SSE2, so vectorization was available on 64bit x86 systems from day one. I think that's at least 17 years if not 20. GCC will use it by default at O2 starting with their next release in about a month. Intel has contributed for a long time but I dont know that this stupidity could be pinned on them. It wouldn't surprise me though.
- pcwalton 4y ago> RMS was right that compilers should be GPL licensed to prevent exactly this kind of thing (and worse things which are haven't happened yet). The problem with this is that it wouldn't solve the problem in question: Intel would just have stuck with their old compiler backend instead of LLVM. Besides, LLVM wouldn't have gotten investment to begin with if it were GPL licensed, since the entire reason for Apple's investment in LLVM is that it wasn't GPL. Ultimately, LLVM itself is a counterexample to RMS's theory that keeping compilers GPL can force organizations to do things: given deep enough pockets, a company can overcome that by developing non-GPL competitors.
- marginalia_nu 4y agoI'm getting flashbacks to the AARD code and Microsoft's attempts to sabotage DR-DOS.
- bee_rider 4y agoHas anyone tried a recent version of MKL on AMD? I assume they were shunting AMD off into an AVX codepath because pre-Zen AMD lacked AVX2 (well, Excavator had I guess...). If they are sending Zen down the generic AVX2 codepaths by default and those are competitive with, say, openBLAS, that seems reasonable, right? Hopefully BLIS will save us all from this kind of confusion eventually.
- penguin_booze 4y agoThis sounds to me very much like VW cheat devices: detect the current situation, and "act accordingly".
- JoshTriplett 4y agoIf Intel had shipped a library/compiler that did just use feature flags and didn't check the CPU vendor, and the resulting code used features that on AMD ran much more slowly than the equivalent unoptimized code, would people blame AMD for the slow instructions, or blame Intel for releasing a library/compiler that they didn't optimize for their competitor's processor? This isn't a hypothetical; quoting https://en.wikipedia.org/wiki/X86_Bit_manipulation_instruction_set https://en.wikipedia.org/wiki/X86_Bit_manipulation_instructi... : > AMD processors before Zen 3[11] that implement PDEP and PEXT do so in microcode, with a latency of 18 cycles rather than a single cycle. As a result it is often faster to use other instructions on these processors. There's no feature flag for "technically supported, but slow, don't use it"; you have to check the CPU model for that. All that said, the right fix here would have been to release this as Open Source, and then people could contribute optimizations for many different processors. But that would have required a decision to rely on winning in hardware quality, rather than sometimes squeezing out a "win" via software even in generations where the hardware quality isn't as good as the competition.
- deleted 4y ago[deleted]
- cbozeman 4y agoWhen it comes to Intel, I am, and have been, so disgusted that they held back computing by about 6-10 years by consistently shipping overpriced, barely improved-upon quad-core processors that I: 1. Put nothing shitty past them. 2. Will never ever purchase their products again. The real problem is the endless pursuit of profit though, instead of the pursuit of ever-advancing, ever-improving technological superiority, and sadly AMD isn't any better in this area I've come to see. The moment they conclusively, provably became better than Intel, they jacked up their price, even though their processors were using the same 7nm process that, at that point, was extremely reliable and had a 93% usable chip ratio. So it turns out as soon as one company gains superiority they immediately become shitbags focused on money instead of focused on the advancement of technology and mankind. It puts anyone with a moralistic stance on what technology should be and how it should be implemented and distributed into a real pickle. I was hoping that AMD would be the better company here, especially given they nearly died, but turns out they also are ready and willing to squander the goodwill of those of us who bought their chips not just when they were on the last legs, but also during their recovery period.
- dang 4y agoRelated: Intel's “cripple AMD” function (2019) - https://news.ycombinator.com/item?id=24307596 https://news.ycombinator.com/item?id=24307596 - Aug 2020 (104 comments) Intel's “Cripple AMD” Function - https://news.ycombinator.com/item?id=21709884 https://news.ycombinator.com/item?id=21709884 - Dec 2019 (10 comments) Intel's "cripple AMD" function (2009) - https://news.ycombinator.com/item?id=7091064 https://news.ycombinator.com/item?id=7091064 - Jan 2014 (124 comments) Intel's "cripple AMD" function - https://news.ycombinator.com/item?id=1028795 https://news.ycombinator.com/item?id=1028795 - Jan 2010 (80 comments)
- snvzz 4y agoRISC is less amenable to this category of BS. Looking forward to RISC-V pushing x86 into retrocomputing territory.