3 ms·
That's one of the more useful ones! It's effectively a low-precision 4-component dot product feeding into an accumulator, which means it is a building block fo
by ack_complete 5y ago
That's one of the more useful ones!
It's effectively a low-precision 4-component dot product feeding into an accumulator, which means it is a building block for larger dot products. Large dot products are very useful both in signal processing for FIR filters, as well as machine learning algorithms. The masking is just a bonus available on most AVX-512 operations and lets you do branchless if conditions.
The majority of vectorized routines I've written have used a multiply-add building block like this, including image resizing, audio resampling and low/high pass filtering, and audio/video compression.
- junon 5y agoHuh, TIL. Would you say that most AVX-512 instructions are then useful in such applications? Given Intel's history of inventing less than useful things (segmented memory, for example) I figured AVX-512 was mostly useless.
- ack_complete 5y agoWell, segmented memory was a pain in the butt but was useful at the time -- not to mention _way_ less of a pain than bank-switched memory. The applications that I mentioned wouldn't use all of AVX-512, but once you're familiar with vectorization it's not hard to tell what they are meant to be used for. CPU designers don't spend silicon on instructions that don't have a use, and AVX-512 does a lot to round out the vector instruction set to make it orthogonal and have less special cases. The bad cases tend to be from instructions that are either too slow or are side effects of the general design. The horizontal add instructions in SSE4, for instance, would have been useful except that they were just as slow as manually doing the shuffles and adds yourself. AVX/AVX2 extended a bunch of SSE2-4 instructions to 256-bit by replicating them across lanes, which led to borderline useless forms like PALIGNR shifting within each 128-bit lane instead of across the entire vector. AVX-512 looks pretty good here though I have seen data about some mask operations being slow, not to mention the whole "slows down the whole chip" issue. That having been said, it's hard to argue that AVX-512 _isn't_ mostly useless, if for no other reason than it being mostly unavailable. Baseline AVX-512 has a market penetration of only 5.6% in the Steam hardware survey and Intel shipping their latest chip without officially supporting it isn't going to help that. Only niche software can afford to use it right now.
- junon 5y agoThanks for the info, this was insightful. :)