3 ms·
I've ordered a tiger lake laptop which should arrive around the end of the month, so I'll be able to test then. So just speculation for now: I'd think AVX512 w
by celrod 6y ago
I've ordered a tiger lake laptop which should arrive around the end of the month, so I'll be able to test then. So just speculation for now:
I'd think AVX512 would still be advantageous for gemm kernels. If you use a 16 x 14 microkernel instead of an 8 x 6 microkernel, you'll reduce memory bandwidth by halving the passes over the cache-sized blocks in memory.
Well optimized implementations don't really have a problem with this (getting close to the CPU's peak flops on AVX2), but it still seems better than not doing it, and is an advantage for code with register tiling more generally.
I suspect that it'd help people using Tullio.jl in Julia, for example.
- gnufx 6y agoYes, if you have two AVX512 FMA units, you use them for GEMM; with a current microarchitecture implementation, if you only have one, you don't. https://github.com/flame/blis/blob/2d8ec164e7ae4f0c461c27309dc1f5d1966eb003/frame/base/bli_cpuid.c#L173 https://github.com/flame/blis/blob/2d8ec164e7ae4f0c461c27309... GEMM is one of relatively few computations with sufficient computational intensity.
- adrian_b 6y agoYes, even on a computer where the throughput of AVX-512 is the same as the throughput of AVX/AVX2, like Tiger Lake, it is usually much easier to reach that throughput with AVX-512 than with AVX/AVX2. The mask registers can eliminate special prologue or epilogue code for the loops and having more and larger registers make it much easier to overlap enough computations so that the latencies of the operations are hidden. I would never buy again an Intel CPU without AVX-512, because that has already for some time been their main advantage and now it remains their only advantage over AMD. Unfortunately for Intel, even if AVX-512 is a great improvement, it cannot compensate for a number of cores half of the competition. Intel needs to launch the 8-core Tiger Lake H no later than March 2021, but it would be better for them if they could do it earlier.
- gnufx 6y ago> it is usually much easier to reach that throughput with AVX-512 than with AVX/AVX2 That's not what I understood from people who've gone through the exercise for GEMM. One probably relevant measurement: on BLIS' generic C GEMM micro-kernel, GCC doesn't get nearly as close to the hand-coded avx512 version as for avx2 with appropriate bock sizes.