10 ms·
I agree with you but would somebody care to enumerate all the ways in which switching is difficult? All i can think is * Intel amt (all enterprises use) * a
by vtesucks 8y ago
I agree with you but would somebody care to enumerate all the ways in which switching is difficult?
All i can think is
* Intel amt (all enterprises use)
* avx-512 (nobody uses)
- davrosthedalek 8y agoIn the HPC segment, there might also be compiler issues.
- ams6110 8y agoThat's true, universities for example will often use Intel compilers on academic licenses, other sites may use other performance-oriented commercial compilers such as PGI.
- greglindahl 8y agoHPC compilers have supported AMD pretty well for the past 15 years, from back when Opteron was the best x64 for a couple of generations.
- bayindirh 8y agoMove aside, I'm an HPC administrator! :) In the HPC world, things are not as clear cut benchmarks, or the vendors' own marketing materials/numbers. First of all, the application you're running may be developed for a specific compiler, and the code sometimes depends on optimization behavior of a compiler. So, changing compilers changes a lot of things. This is why we have both Intel's tools, and GCC toolchain fully supported. For example, LAPACK and its siblings take compiler behavior and CPU specifications into consideration while compiling in an optimized way to maximize its performance IIRC. Also, there's no guarantee that Intel's compilers are fastest on Intel hardware. In the days of Opteron 6100s, using Intel compilers, we were able to beat Intel processors of the same era. You heard it right: Compile using intel compiler with specific flags, run on AMD CPUs, get higher performance, profit! Intel's AVX512 is well used and abused in HPC world, however AMD's HPC performance is not as bad as jandrewrogers implied in his comment [0]. AMD is originally an FPU company, and while their scalar instructions may lack on paper, they run really fast. In HPC world, the CPU/board architecture becomes irrelevant after some point. SpecCPU benchmarks are the ultimate benchmarks, because their behavior is compiler agnostic and push every aspect of the CPU very very hard. If you can get the same SpecFP with an Intel part, you can get the more or the less same performance on real workloads. If you have any other questions, you can AMA. I'll try my best to answer. Funny addenda: We have some applications used widely by users, and when fully optimized, some older Intel CPUs outpace the newer ones by a significant margin. This is some heavy handed, exotic optimization. [0]: https://news.ycombinator.com/item?id=18628592 https://news.ycombinator.com/item?id=18628592
- jcranmer 8y ago> Intel's AVX512 is well used and abused in HPC world, however AMD's HPC performance is not as bad as jandrewrogers implied in his comment [0]. I know earlier AMD processors didn't actually have a 256-bit support, so AVX instructions were actually implemented by soaking up two 128-bit lanes (it helps that AVX doesn't have many instructions that actually permit you to move data between the two 128-bit slices of a 256-bit vector). For their AVX-512 performance to not be absolutely horrible, I take it they've actually built real AVX-512 units at some point?
- bayindirh 8y agoFrom what I've read now [0], it looks like AMD still uses 2 x 128bit AVX units to execute AVX2 instructions. Also, AMD is always coming a generation behind Intel in terms of FP instructions sets, so Zen doesn't support AVX512. According to WikiChip [4], Zen 2 actually has 256 bit FPU paths. I was unable to find a credible benchmark for Zen 2, so I can't talk about its performance. However, when analyzed from the perspective I've given below, it's not hard to assume that Zen 2 is a heavy hitter in terms of floating point performance. However, the interesting part is, when you look to SpecCPU 2017 FP Rate [1], AMD Epyc 7601 [2] system has a similar per core performance with a much bigger Intel Xeon Platinum 8180 [3] system. Why interesting? * AMD's per core base (lowest) rate is 4.1875. * Intel's per core base (lowest) rate is 4.3482. * AMD is running GCC compiled code. * Intel is running Intel compiled code. * Intel has higher clock speed. Intel has some CPUs (like Gold 5118, Gold 6148) which have per core base rate of ~5.125. These are the CPUs are considered as HPC processors, and used by a lot of people. As I said before, it looks like Zen 2 is going to be a better HPC processor than Zen. Zen looks like a very good Enterprise processor now. So with my hat, I can conclude that not having 512 bit hardware is not a crippling omission. Addenda: I forgot to say that Intel has something called "AVX frequency". Since AVX, AVX2 and AVX512 has tremendous power requirements when compared to other operations, Intel lowers CPU to an undisclosed frequency. When I last checked, AVX frequencies of Intel CPUs that we use weren't in the technical guides and were not public in any way. So, the peak SpecFP Rate is not very different from the base ones. Also, since the CPUs thermal budget is very constrained during AVXx operations, other ports' speed is also reduced. So at the end of the day, AVX512 is not a free turbo boost in HPC environments and heavy/continuous loads. [0]: https://en.wikichip.org/wiki/amd/microarchitectures/zen#Floating_Point https://en.wikichip.org/wiki/amd/microarchitectures/zen#Floa... [1]: http://spec.org/cpu2017/results/rfp2017.html http://spec.org/cpu2017/results/rfp2017.html [2]: http://spec.org/cpu2017/results/res2018q4/cpu2017-20180917-08862.html http://spec.org/cpu2017/results/res2018q4/cpu2017-20180917-0... [3]: http://spec.org/cpu2017/results/res2017q4/cpu2017-20171017-00119.html http://spec.org/cpu2017/results/res2017q4/cpu2017-20171017-0... [4]: https://en.wikichip.org/wiki/amd/microarchitectures/zen_2#Key_changes_from_Zen.2B https://en.wikichip.org/wiki/amd/microarchitectures/zen_2#Ke...
- simias 8y agoGiven that I mostly work on embedded software I'm not the best person to ask but I suppose that if you're writing software that uses heavy performance optimizations going from one CPU model to an other is not necessarily trivial. On top of that most server farms wouldn't completely ditch their old computers to replace them in one go, so you end up with two different vendors to deal with, potentially two versions of your software etc... That increases the maintenance burden quite significantly.
- spamizbad 8y agoIntel might be giving major buyers sweetheart deals on their chip prices. So while an EPYC 7371 might get discounted 20% to an Amazon, Intel may be discounting 50-70% (Gain of salt: just speculation on my part). And when you consider these CPUs are in systems where most of your costs are likely in RAM, NVMe, and supporting infrastructure. These are high margin parts, so I imagine there's wiggle room for volume.
- selectodude 8y agoNot a chance. Intel is selling every CPU they can print. Enterprise customers are paying a premium at the moment.
- JackFaker 8y agoSTH previously reported that Intel began offering Xeon discounts to smaller organizations that asked for AMD EPYC price quotes. https://www.servethehome.com/intel-is-serving-major-xeon-discounts-to-combat-amd-epyc/ https://www.servethehome.com/intel-is-serving-major-xeon-dis...
- selectodude 8y agoThat's a niche market. Intel has the larger players by the balls. Very few companies are comparison shopping AMD vs Intel.
- jburgess777 8y agoYou say that, but AWS has still launched EC2 instances with both AMD and ARM CPU cores recently. I imagine Amazon buy enough Intel CPUs for them to take notice.
- Nexxxeh 8y agoI wonder if the primary driver for buying and offering alternative platforms for EC2 instances was to send Intel a message. "We aren't afraid to go to AMD if you don't make your offering more competitive."
- jasode 8y ago>avx-512 (nobody uses) An example of people using those cpu instructions would be buyers of Intel's proprietary C++/Fortran compiler.[1] The reason companies pay ~$1700 license instead of using free compilers such as GCC and Clang is to specifically take advantage of the latest advanced Intel cpu instructions. Example buyers of Intel's C++ compiler would include high-frequency trading firms and HPC labs. I wouldn't be surprised if Google, Facebook, and Amazon also bought Intel Parallel Studio compiler licenses for some of their workloads. [1] https://software.intel.com/en-us/parallel-studio-xe/support/license https://software.intel.com/en-us/parallel-studio-xe/support/...
- deleted 8y ago[deleted]
- celrod 8y agoRecent LLVM and GCC make good use of avx512. Although for GCC, you need to use the option -mprefervectorwidth=512. Unfortunately benchmarks on websites like Phoronix do not make use of them. But when it comes to numerical computing, it is a boon. Much easier to take advantage of than the GPU. I run (Monte Carlo) simulations that take hours or days. These can be vectorized, but I've never heard of someone being able to run them on a GPU. However, folks bring graphics cards up every time I mention (my love of) avx512. There is always a first, so I do really want to find the time to play around with it, and see how many mid-sized chunks can be woven together. And how memory/cache plays out when breaking things into small pieces. The last Monte Carlo simulation I ran took a few days to get 100 iterations. The MC iterations themselves were chains of Markov Chain Monte Carlo iterations. Each of these MCMC iterations takes several seconds. Therefore, to move to a GPU, I'd like to parallelize between MC iterations, and also within the MCMC iterations. On a CPU all you have to do is vectorize the MCMC iterations, and then run the chains in parallel.
- derefr 8y ago> I've never heard of someone being able to run them on a GPU I'm surprised; intuitively (though, mind you, as someone who has never done GPGPU programming, only read articles about it), I'd think some combination of 1. a CPU-RNG-seeded simplex-noise kernel, for per-core randomness; and 2. a cellular-automata kernel embedding of your simulation logic, would let you do MC just fine.
- jandrewrogers 8y agoThere are a couple more reasons, all related to loss of optimization that can offset any nominal price-performance gains. Counterintuitively, people who are the most sensitive to performance often have the most to lose. AMD implements some important scalar instruction set extensions as microcode, not in silicon, so if you have an application that uses them heavily (and some of these instructions are significant optimizations over generic C code) you will see a drop-off in performance. Highly optimized/efficient code for Intel microarchitectures become a lot less so on the significantly different AMD microarchitecture. The effects are not small and re-optimizing for a different microarchitecture can be a lot of work depending on the application.
- BeeOnRope 8y ago> AMD implements some important scalar instruction set extensions as microcode, not in silicon Do you have any examples other than pdep and pext? Although these happen to be my two favorite scalar instructions, I would hesitate to call them important. Compilers won't just generate these from normal source [1], and I would call their use extremely niche at the moment (things like chess engines, I'm looking at you). They aren't even available on Intel Ivy Bridge and Sandy Bridge machines, which still make up a big enough fraction of data center machines. So I'm pretty sure the number of entities avoiding switching to AMD because of heavy pdep and pext use is pretty close to zero. Maybe you have some other instructions in mind though? > Highly optimized/efficient code for Intel microarchitectures become a lot less so on the significantly different AMD microarchitecture. This was somewhat true in the past, and probably hit its peak in the P4 vs Athlon/Opteron era. However, it is pretty much incorrect for Zen. Although the details of the hardware implementation might differ (and unless you are an insider you can mostly only guess at this), as an optimization target for software, Zen is very similar. It has a similar width, similar cache design both for data and instructions, similar instruction latencies and throughput, and so on. In fact something like Zen is as similar to Haswell as Haswell is to say Ivy Bridge. The primary exception is AVX/AVX2 code, where Zen implements everything internally as 128-bit operations. In this area you might make some different decisions if targeting Zen - but the gap is not huge. --- [1] What I mean is they won't generate them any scenario other than directly calling the x86-specific builtin/intrinsic for that exact instruction.
- secabeen 8y agoThis isn't a huge one, but we're a VMWare shop, and you can only VMotion (move live machines) VMs between server on the same chip class or lower. When we had a mix of AMD and Intel servers, we had to put them in different clusters, and could only migrate between clusters when the VM was powered-off.
- wstuartcl 8y agoYou can, however simply power down the vm and then use vmotion to migrate across from intel to amd. Not "live" but pretty close to it when migrating to the new cpu cluster.
- secabeen 8y agoYes, but depending on your architecture, that may result in a service outage or degradation. It's really nice to be able to VMotion live machines without any impact to users. Sure, Devops, durable machines, chaos monkey, etc. etc.; we don't develop all of the code our users run, and not all of it has those features yet, or will ever.
- btian 8y ago> * Intel amt (all enterprises use) Is that true? I work for a hundred billion $ tech company, and Intel AMT is disabled fleet wide due to security issues.
- Teknoman117 8y agoI work at one of those as well. It's disabled on all our machines to.
- ascar 8y agoOdds aren't that bad that you work for the same one :D
- justapassenger 8y agoTech companies would tend to disable them. Non-tech enterprises like things like that.
- loeg 8y agoavx-256 performance per core is half of Intel's, and people use that.
- BeeOnRope 8y agoIt depends. AMD Zen has 4 128-bit SIMD units, while Skylake-S (and earlier) Intel chips have 3 256-bit SIMD units, and Skylake-X/SP (hereafter SKX) chips additionally have 1 or 2 512-bit SIMD units (which overlap with the 256-bit units). Now not all units can run all instructions. E.g., Intel chips run FP instructions on only 2 of those units, but AMD can run FP instructions on all 4, so in that sense they are even. However, not all FP instructions can run on all units on AMD: FP multiplications run on only 2 units, and FP additions run on two units - so if you are doing all multiplies, both AMD and Intel can do 2 per cycle and since Intel is twice (quadruple on SKX with 2 FMA units) as wide, then Intel is twice as fast. If you do a 1:1 mix of mul and add, however, then AMD and Intel may be tied - but then AMD is further hamstrung by only have 2 128-bit load units, vs Intel's 2 256-bit load units (512-bit on SKX) - so it is entirely possible that many kernels are limited by load/store throughput on AMD. For things like in-lane shuffles, AMD and Intel (pre-SKX) have the same throughput: AMD has 2 128-bit shuffle units, and Intel 1 256-bit one. For cross-lane shuffles the 128-bit AMD units really struggle since multiple ops are required and Intel wins big. So AMD's AVX-256 perf being half of Intel's is more or less the worst case, and many cases will see closer performance. If Zen 2 doubles everything to 256-bits everything will change dramatically.
- paulmd 8y agoAll Skylake-X chips actually have the second 512b SIMD unit, including the i7s. Intel's initial documentation here was incorrect, InstLatX64 determined that both units are enabled. https://twitter.com/InstLatX64/status/935966607188295680 https://twitter.com/InstLatX64/status/935966607188295680
- BeeOnRope 8y agoYes, for now - but I don't think I implied otherwise? I said Skylake-X/SP. Certainly not all Skylake-SP (aka "Scalable Xeon" or whatever Intel calls them) server chips have 2 FMA units. BTW, I believe Intel also documents this fact about X chips on ARK now as well.
- MrBuddyCasino 8y agoIts an open secret that Intel adds custom accelerator hardware to their server chips for the big corps, that are undocumented/disabled for everyone else. I have no idea if AMD does this too.