4 ms·
Looks like they think they're still winning regardless of the price and that simply bumping core count to be the kings and bringing the price back to the Haswel
by slizard 9y ago
Looks like they think they're still winning regardless of the price and that simply bumping core count to be the kings and bringing the price back to the Haswell-EP level high (rather than Broadwell-EP crazy) will be enough.
What also shows that they seem to be confident is that they're further segmenting the market based on the PCIE lane count to push everyone wanting >32 lanes into the >$1k regime.
All in all, the cool thing is not the i9s and high core counts which you could get even before by plugging a Xeon chip into a consumer X99 mobo (though you'd have to pay some $$$), but the new cache hierarchy which will give serious improvements in well-implemented, cache friendly codes!
- slizard 9y ago...and of course AVX-512 for the lucky ones that can get significant benefit from such a wide SIMD (also considering the very likely significant clock limit for AVX instruction streams).
- marmaduke 9y agoeven chips with AVX2 on all cores slow down when it's fully used. The Xeon Phi has a pretty low clock, 1.3 ghz iirc. Still, it gets you GPU style performance on vector workloads without needing separate hardware and software stack.
- valarauca1 9y agoeven chips with AVX2 on all cores slow down when it's fully used Not really. Xeon Phi's clock low because the die is massive. The downclocking for AVX started with Knights Landing. My Boardwell-EP Xeon stays at 3.0Ghz even when I (ab)use AVX2.
- AbacusAvenger 9y agoI tried AVX512 on a Xeon (non-Phi) part recently and it was extremely underwhelming. The workload (OpenMP-parallelized n-body) was actually slower with AVX512. Since it was under virtualization and I didn't have access to the bare metal hardware or to performance counters, I have no way of knowing why, but I'm almost certain it was because it lost all-cores Turbo and downclocked aggressively. It had previously scaled almost linearly going from SSE to AVX/AVX2, but it regressed with AVX512.
- smitherfield 9y agoIt might be your processor only supported AVX512 in emulation — the article makes it sound like only the Phi currently supports it natively.
- AbacusAvenger 9y agoSo they implemented AVX512 on the Xeon server parts in microcode? That seems crazy.
- smitherfield 9y agoIt's fairly common practice with bleeding-edge vector instructions. The reasoning (assuming it is the case here) is that a theoretically-minor performance regression (the cost of converting 1x AVX512 to 2x AVX2 in microcode) is usually much preferred over a CPU exception when attempting to run a binary with AVX512 instructions on a server. It also means you don't need a $15000 chip to test your AVX512 code.
- marmaduke 9y agoThat's disappointing. It may be a bad interaction between OpenMP and AVX512 (cache pressure etc)? I've also seen reliable increases in performance up through AVX2 but when I tried to run same code on a Xeon Phi, it fell short of the plain Xeon.
- slizard 9y agoWhat scaling are you referring to when talking about "linear"? Did you really get 2x absolute performance going from 128-bit to 256-bit SIMD (regardless of the uarch, e.g. both with SBE and HSW with the former having a relatively poor cache performance). I'd be surprised, but if your code is in the >>O(10) flops/byte regime [1] and especially friendly instruction stream too. Otherwise, I'm skeptical. Putting aside the "scaled almost linearly" statement, I'm not surprised that AVX-512 did not give the expected benefits ootb, you suddenly need double the amount of data loaded into the registers for every 512-bit instruction. You'd also quite like want to make sure masking is used effectively. [2] [1] http://people.eecs.berkeley.edu/~kubitron/cs258/lectures/lec12-Merrimac.pdf http://people.eecs.berkeley.edu/~kubitron/cs258/lectures/lec... [2] https://software.intel.com/en-us/node/523777 https://software.intel.com/en-us/node/523777
- burntsushi 9y agoThis[1] seems to suggest otherwise? Or am I misinterpreting it? [1] - https://computing.llnl.gov/tutorials/linux_clusters/intelAVXperformanceWhitePaper.pdf https://computing.llnl.gov/tutorials/linux_clusters/intelAVX...
- slizard 9y ago> Xeon Phi's clock low because the die is massive. That's not the main reason. The main reason is perf/W for highly parallel workloads. > The downclocking for AVX started with Knights Landing. You're mixing things up here and that statement is incorrect too. AVX throttling started with Haswell-EP [1,3] (Intel kept it quite hush-hush avoiding mentioning it in products specs and such). Secondly, Xeon Phi is the HPC product family and KNL is the codename of the 2nd generation of these arch [2] > My Boardwell-EP Xeon stays at 3.0Ghz even when I (ab)use AVX2. In that case you're most likely either not using more than 1-2 cores or you're overclocking (or perhaps monitoring incorrectly), see [1,3]. [1] http://images.anandtech.com/doci/8423/AVXTurboHaswEP.png http://images.anandtech.com/doci/8423/AVXTurboHaswEP.png [2] https://en.wikipedia.org/wiki/Xeon_Phi https://en.wikipedia.org/wiki/Xeon_Phi [3] https://www.microway.com/knowledge-center-articles/detailed-specifications-of-the-intel-xeon-e5-2600v4-broadwell-ep-processors/ https://www.microway.com/knowledge-center-articles/detailed-...