5 ms·
Boost your benchmarks with AVX-512 by subscribing to Intel+
by tmccrary55 5y ago
Boost your benchmarks with AVX-512 by subscribing to Intel+
- amluto 5y agoWhat benchmarks? Not a lot of real programs actually use AVX512, and a lot of the ones that did discovered that it made performance worse, so they stopped. (The issue is that AVX512 (the actual 512-bit parts, not the associated EVEX and masking extensions) may well be excellent for long-running vector-math-heavy usage, but the cost of switching AVX512 on and off is extreme, and using it for things like memcpy() and strcmp() is pretty much always a loss except in silly microbenchmarks.) To be clear, I don't like the type of product line differentiation that Intel does, and I think Intel should have supported proper heterogenous ISA systems so that AVX512 and related technologies on client systems would make sense, but I don't think any of this is nefarious.
- hajile 5y ago> What benchmarks? Not a lot of real programs actually use AVX512, and a lot of the ones that did discovered that it made performance worse, so they stopped. This paints a picture of what happened to Intel's fab process rather than AVX-512 itself. AVX-512 was proposed back in 2013. It was NOT designed for desktops. The original designs were for their Phi chips (basically turning a bunch of x86 cores into a GPU). These Phi chips ran between 1 and 1.5GHz, so power consumption and clocks were always matched up. Intel wanted to move these instructions into their HPC CPUs. The problem at hand was ultra-high turbo speeds. These speeds work because the heat is a bit spread out on the chip and a lot of pieces are disabled at any given time. With AVX-512, they had 30% of the core going wide open for the vector units (not to mention the added usage from saturating the load/store bandwidth). They wanted 10nm and then 7nm to fix these issues. At their predicted schedule, 10nm would have launched in 2015 and 7nm in 2017. Given the introduction of AVX-512 in 2013, they had plenty of time. In fact, the first official product was Knight's landing in 2016. Skylake with AVX-512 didn't launch until 2017 when they were supposed to be on 7nm. Intel was forced to backport their designs to 14nm++++++++++++ which forced a deal with the devil. They had to downclock the CPU to keep within thermal limits, but this slowed down EVERYTHING. Maybe they could have created a separate power plane for AVX, but this would be a BIG change (and probably politically infeasible). What happens with the downclocking? If you run dedicated AVX-512 loads, then maybe you should have been looking at Phi instead. If not, mixed loads suffered overall because of the lower clockspeeds. Their second revision of 10nm superfin is still a generation larger (well, probably a half-generation) than what they anticipated. There might still be downclocking, but I'd guess that it's nowhere near what previous iterations required. TL;DR -- AVX-512 was launched two nodes too early which screwed over performance. It should become acceptable either with 10nm Superfin or 7nm when it launches so the CPU doesn't have to downclock constantly.
- azalemeth 5y agoIt's worth saying that in my life as and academic, the only program I've ever used that has commonly benefited from avx-512 in my work is Gromacs. Gromacs is a molecular dynamics program that creates beautifully graphical simulations and often appears as-is in benchmarks (the so-called ns-per-day metric). Although almost indescribably complex it's also fundamentally strangely straightforward in what it does, and avx-512 does indeed make it significantly faster. The overwhelming bulk of my research does not use Gromacs, however. Mkl, matrix algebra and similar hybrid MPI tasks, yes, but oddly there the extensions really don't seem to do much, and frankly the inferior memory architecture of the Xeon platinum nodes we use makes itself apparent. I frequently get annoyed that Intel charge you a fortune for CPUs into which you can put a ton of ram, and then marketing makes the CPU's exclusive feature an instruction that doesn't seem to help that much in my actual workloads. They really should just shunt it to some specialised product and use the die space for something else. I'm with Linus here.