6 ms·
Question is, how many cores did we have to sacrifice to get AVX512? Intel moved from shipping 10 core CPUs to 8 core CPUs with AVX512. Clock speed may not go do
by NohatCoder 5y ago
Question is, how many cores did we have to sacrifice to get AVX512? Intel moved from shipping 10 core CPUs to 8 core CPUs with AVX512. Clock speed may not go down much, but the price of that is a very high power consumption on AVX512 workloads.
As best I can tell, it would be reasonable to expect around 12 AVX2 cores operating on the same power budget as the current 8 AVX512 cores. Does the average user have enough AVX512 workloads to make that tradeoff worth it?
- zigzag312 5y agoFor SIMD workload performance increase of AVX512 over AVX2 is 2x. How many cores would you need to add to double the performance of 8-core CPU? Any media processing can generally gain significant performance with SIMD instructions. Even web browsing is a workload that is affected, as (among other things) libjpeg-turbo uses SIMD instructions to accelerate decoding. It doesn't yet use AVX512, but it does use AVX2. If there are gains with AVX2, why wouldn't there be with AVX512. Even if AVX512 would remain 256 bits, it would be an improvement over AVX2 as new instructions enable acceleration of workload that were not possible before, plus there is an increase in productivity/ease of use. However, AVX512 won't be used much for consumer applications until there is big enough market penetration of CPUs that support these instructions, as it has been the case with any new SIMD instructions when they were first introduced.
- dragontamer 5y ago> Any media processing can generally gain significant performance with SIMD instructions To a limit. JPEG (and many video codecs based on JPEG) have 8x8 macroblocks, which means the "easiest" SIMD-parallel is 64-way. And AVX512 taken 8-bits at a time is in fact, 64-way SIMD. To get further parallel processing after that, you'll probably have to change the format. GPUs go up to 1024-way NVidia blocks (or AMD Thread groups), which are basically SIMD-units ganged together so that thread-barrier instructions can keep them in sync better. 1024-work items corresponds to a 32x32 pixel working area. But that's no longer the format of JPEG. It'd have to be some future codec. Maybe modern codecs are seeing the writing on the wall and are increasing macroblock size for better parallel processing 10 years into the future (they are a surprisingly forward looking group in general).
- janwas 5y ago> Maybe modern codecs are seeing the writing on the wall and are increasing macroblock size for better parallel processing 10 years into the future We did indeed do this for JPEG XL - the future is now :) 256x256 pixel groups are independently decodable (multi-core), each with >= 64-item (float) SIMD.
- cma 5y agoAVX-512 lines up with 64-byte cache lines, it seems like it would be a huge change to go bigger.
- dragontamer 5y agoNVidia GPUs are 32 wide warps, AMD CDNA are 64 wide. That's 1024 bit and 2048 bit respectively. Cache lines are probably 64 wide for the purpose of burst length 8 (64 bit burst length 8 is 64 bytes / 512 bits).
- Tostino 5y agoI'm confused by your argument that this is similar to prior rollouts of new instruction sets. In the past when Intel pushed into instruction set, it generally went to their whole line of chips very quickly. That's not what I've seen with AVX at all.
- zigzag312 5y agoI was referring to software adoption. Software adoption cycle is similar, as wider software adoption happens only when there is enough market penetration of CPUs that support new instructions. There was AMD's 3DNow! that saw limited software adoption, because Intel didn't support it. Newer instruction sets are getting adopted progressively slower as consumers are replacing computer less often, AMD is slower at adopting each new AVX instruction set and Intel is getting more aggressive with market segmentation. Because market penetration of new instruction sets is getting slower, SW adoption is also much slower.
- janwas 5y agoTotally agree software uptake is the pain point. Not even all gamer CPUs have SSE4 (https://store.steampowered.com/hwsurvey https://store.steampowered.com/hwsurvey), so it seems that runtime dispatch is unavoidable. Given that, if we can afford to generate code for new instruction sets and bundle it all into one slightly larger binary/library, the problem is solved, right? Highway makes it much easier to do that - no need to rewrite code for each new instruction set. As to Intel's market segmentation, we target 'clusters' of features, e.g. Haswell-like (AVX2, BMI2); Skylake (AVX-512 F/BW/DQ/VL), and Icelake (VNNI, VBMI2, VAES etc) instead of all the possible combinations.
- adwn 5y ago> For SIMD workload performance increase of AVX512 over AVX2 is 2x It's less than 2x, because the core downclocking for AVX-512 on older CPUs is higher than for AVX2: 60% vs 85% on Skylake, so only a ~1.4x speedup. Newer CPU architectures do not downclock, though.
- zigzag312 5y agoTo be fair, multicore workload also causes CPU to operate at lower frequency than at single core workload. In the end, multi-threading and SIMD are both types of parallel processing and each has pros and cons.
- account42 5y ago> Newer CPU architectures do not downclock, though. Newer CPU architectures may not have enforced downclocks, but they will eventually respond to the higher thermal output.
- brigade 5y agoAmdahl's law. JPEG decoding especially - the majority of the decode time these days is entropy decoding that cannot be parallelized by either SIMD or threads.
- brigade 5y agoSunny/Cypress Cove had lots of other area growth than just AVX-512 [1]. The better comparison is Skylake vs Skylake-SP, from which AVX512 is estimated to cost about 5% of the tile area [2]. So that’s the area cost of about 2 cores in the 28 core chip. [1] https://en.wikichip.org/wiki/intel/microarchitectures/sunny_cove#Key_changes_from_Palm_Cove.2FSkylake https://en.wikichip.org/wiki/intel/microarchitectures/sunny_... [2] https://www.realworldtech.com/forum/?threadid=193291&curpostid=193291 https://www.realworldtech.com/forum/?threadid=193291&curpost...
- wtallis 5y ago> The better comparison is Skylake vs Skylake-SP, from which AVX512 is estimated to cost about 5% of the tile area That analysis might be slightly underestimating the area penalty of AVX512, because the consumer Skylake cores that didn't have AVX512 execution units still reserved space for the AVX512 register file. (And given that fact, it's all the more surprising that while Intel was repeatedly refreshing 14nm Skylake for the consumer market, they never added the rest of the AVX512 bits or redid the layout of the consumer cores to reclaim the blank space of the register file.)
- brigade 5y agoUnlike the ALUs, it's trivial to locate and determine the area of a 512 bit x 168 register file. No way would Kanter have missed that in his analysis.
- wtallis 5y agoIt's not really a matter of whether Kanter missed those spots, but of whether subtracting them from the total core area is relevant to the question he was trying to answer. (Also, both the ALUs and the register files are trivial to locate by comparing die photos of consumer and server versions of the Skylake core.) If you're going to hypothesize about a rearranged Skylake core that no longer reserves space for the full 512 bit vectors in the register file, then you probably should also try to estimate the other area savings that would result from truly removing AVX512 from the core design at a high level, rather than merely masking off a few regions of silicon that serve no purpose other than AVX512 support. And answering that question is a lot more subjective than simply tallying up the area occupied by the extra execution units hanging off the side of the AVX512-enabled server Skylake cores. Or, to put it another way: there unquestionably is an area cost to AVX512 beyond that of the extra execution units tacked on to SKL-X cores. But that cost was already being paid by consumer CPUs years before SKL-X/SKL-SP shipped. So it wasn't really a contributing factor to the loss of two cores when Skylaked-derived Comet Lake was succeeded by Rocket Lake.