3 ms·
I work with GPUs professionally, but I've written a lot of AVX-512 code (e.g., https://NN-512.com https://NN-512.com) and AVX-512 is a big step forward from AVX
by 37ef_ced3 5y ago
I work with GPUs professionally, but I've written a lot of AVX-512 code (e.g., https://NN-512.com https://NN-512.com) and AVX-512 is a big step forward from AVX2.
The regularity and simplicity and completeness of the instruction set is a big win.
The lane predication (masking) of every instruction is useful when the data doesn't fit the vectors (and in loop tails) but it has numerous other uses, too. For example, it makes blending (vector splicing) and partial operations easy.
The flexible permute instructions (two data vectors in, one control vector in, and one data vector out) are fast and enormously useful. Anyone who has puzzled with AVX2 will breathe a sigh of relief.
The register file is big enough (32 vectors, each 16 floats wide) that register starvation typically isn't an issue. For example, you don't worry about having registers for your permutation control vectors. And the predicate masks are in another set of registers, which again are plentiful (eight!).
The easing of alignment requirements (unaligned load/store instructions operating on aligned addresses are as fast as aligned loads/stores) is also a big win.
AVX-512 is a real pleasure to use.
- dragontamer 5y agoAvx512 probably needs bpermute for parity with GPU shared memory. Arbitrary swizzles both in the forward permute (currently in AVX512) and backwards direction (GPU only) is as necessary and proper as gather + scatter. Any program that uses one is highly likely to use the other. -------- Butterfly permutes should be especially accelerated, as that pattern continues to show up. It seems like arbitrary permutes / bpermutes are expensive to implement in hardware, but butterfly permutes are the fundamental building block. Butterfly permutes are needed from FFT to scan/prefix sum operations. It's also fundamental (and simpler at the hardware level than arbitrary bpermute/permute) AMD implements the butterfly permute in DPP instructions. NVidia provides a simple shfl.bfly PTX instruction. Butterfly networks (and inverse butterfly) are how pdep and pext are implemented under the hood. http://palms.ee.princeton.edu/PALMSopen/hilewitz06FastBitCompression.pdf http://palms.ee.princeton.edu/PALMSopen/hilewitz06FastBitCom... ----------- IIRC,the butterfly / inverse butterfly network can implement any arbitrary permute in just log2(n) steps.
- psykotic 5y agoI think you need 2 log2(n) - 1 butterfly stages to realize an arbitrary permutation.
- wyldfire 5y agoI work on some tooling for Hexagon DSPs and I wonder how the HVX instructions compare against AVX-512 (for integers at least). Different targets / use cases but I'm curious how it stacks up.