4 ms·
I wonder why the cascaded triple for loop for GEMM is not just hardcoded in assembly for each architecture. And while we are at it, why is MKL > 5GB again?
by WanderPanda 3y ago
I wonder why the cascaded triple for loop for GEMM is not just hardcoded in assembly for each architecture. And while we are at it, why is MKL > 5GB again?
- tomrod 3y agoDoes this address your MKL question?[1] If so, sounds like we're two of today's lucky 10,000![0] Although for this esoterica, I'd probably reduce that to 10. [0] https://xkcd.com/1053/ https://xkcd.com/1053/ [1] https://community.intel.com/t5/Intel-oneAPI-Math-Kernel-Library/Size-of-MKL-libraries/td-p/1090402 https://community.intel.com/t5/Intel-oneAPI-Math-Kernel-Libr...
- zzzoom 3y agoMost decent BLAS implementations use some variation of the kernels in K. Goto et al. "Anatomy of High-Performance Matrix Multiplication" written in assembly.