4 ms·
Intel defends AVX-512 against critics who wish it to die a 'painful death'
- formerly_proven 6y ago"Intel defends fragmented ISA obviously meant to put market segmentation into binaries just like they always wanted to in the good ol' Itanium days by doing their usual 'up to 300X performance increase with this simple trick!' schtick" Also let's not gloss over the hilarity of trying to compete against "literally a grid of ALUs connected to memory" with x86 cores tuned for per-thread performance for a delightfully parallel task such as "multiple these two billion tensors". HPC as a market at least makes some sense, since HPC applications feast on double precision FLOPS and largely don't care about SP, the raison d'etre of GPUs.
- captainbland 6y agoI find this argument that AVX-512 is only really useful for HPC kind of odd. Isn't it going to be useful in all the same kinds of domains that are traditionally quite SIMD friendly like video encoding, image manipulation and such? There must be a lot of professionals doing these kinds of tasks on their laptops and not necessarily wanting to add a discrete GPU or going to the cloud to get reasonable performance.
- formerly_proven 6y agoThere is no way to do video or image processing on a CPU with reasonable performance. But if you really wanted to, using the IGP would make more sense, since they usually have around 10x-20x the throughput of all the CPU cores next to them, so instead of being choked by both the low memory bandwidth and the low throughput, you only have to contend with the low memory bandwidth. All the stuff x86 burns so many transistors and so much power on just straight up doesn't matter for DSP. And, instead of trying to make these cores somewhat less inept at a task they are not competitive at anyway, you could instead spend the transistors on the IGP or cache or interconnect.
- devwastaken 6y agoPopular video encoders heavily target modern cpu instructions, because that is where a great number of people will be encoding their own videos from. GPU encoding is faster if it has special hardware support, and other kinds of hardware could do it faster, but the algorithms for it leave out a number of quality options only available on the cpu versions. For example, hevc with AVX2.
- formerly_proven 6y agoPeople who care about squeezing the most out of available bitrate use things like x264, but that's fewer people than you might think. And of course open source aficionados will use CPU encoders, since no hardware-assisted encoders are open source (afaik. Maybe AMD's open source driver can drive their video codecs, I dunno, AMD cards are irrelevant when it comes to content production, except when they're stuck in a Mac and use Metal). In content creation, creating deliverables for e.g. YouTube or other video platforms it makes little sense to waste time on CPU encoding, when Intel QuickSync or nvenc are an integer factor faster and give you essentially the same quality, at a higher bitrate, which doesn't matter anyway, because the platform is re-encoding it 100 % of the time.
- ohazi 6y agoIt makes more sense when you read Torvalds' initial argument. AVX-512 draws so much power that (especially in laptops) they have to lower the base clock of the entire CPU. This doesn't matter too much if you're doing HPC, because lowering the clock 20% but gaining 200% in throughput is still a net win. But the computer can't change the clock speed instantly, so there's a fairly large penalty for just executing some AVX-512 instructions and then immediately trying to go back to other tasks (which is what happens if you try to use AVX-512 to implement memcpy). Or if you're trying to simultaneously use other parts of that core for other tasks, they'll be slower. So back to whether it'll speed up video encoding... Perhaps, but the original complaint was that "when is it appropriate to use AVX-512" already has a lot of asterisks on a Xeon in a data center... It's going to have even more asterisks in a power constrained laptop, which suggests that bringing it to laptops may not be the best option, given the availability of other options. Apparently it's difficult for Intel to look at AVX-512 objectively. The fact that their HPC customers love it doesn't imply that it's well designed or appropriate for laptops.
- captainbland 6y ago> It makes more sense when you read Torvalds' initial argument. I did but: > AVX-512 draws so much power that (especially in laptops) they have to lower the base clock of the entire CPU. This doesn't matter too much if you're doing HPC, because lowering the clock 20% but gaining 200% in throughput is still a net win. Is still basically true for encoding tasks and other SIMD-heavy operations. > But the computer can't change the clock speed instantly, so there's a fairly large penalty for just executing some AVX-512 instructions and then immediately trying to go back to other tasks (which is what happens if you try to use AVX-512 to implement memcpy). This is interesting on this point: https://travisdowns.github.io/blog/2020/01/17/avxfreq1.html https://travisdowns.github.io/blog/2020/01/17/avxfreq1.html There is a cost to switching frequencies but the cost is generally lower than the performance impact of the new frequency itself. Which kind of makes sense given how frequently dynamic frequency scaling and turbo boost/turbo core like to change the frequency of the CPU. That said, just because somebody's decided to try and use a CPU feature sub-optimally (e.g. by implementing memcpy with AVX-512 so that it gets used sporadically in the context of otherwise non-AVX related work), I don't think that necessarily justifies removing it. It does sound like the power management for it ought to be fixed, though. Is it just me or is the idea of having power management for a specific instruction extension rather than integrating it into the rest of the dynamic frequency scaling/turbo boost TDP management kind of a smell? But I do kind of buy into the idea in the sibling comment that the transistors might have been better spent on the iGPU, though - at least in the short term.