3 ms·
Blis is very bad at small sizes, last I checked. Their approach was based on having a single microkernel, and setting up calls to it. Last I checked, blasfeo
by celrod 4y ago
Blis is very bad at small sizes,
last I checked. Their approach was based on having a single microkernel, and setting up calls to it.
Last I checked, blasfeo did not support AVX512, and this performed poorly on CPUs supporting it.
I'm not really familiar with xsmm, but someone showed me this: https://haampie.github.io/smm-bench/cascadelake/ https://haampie.github.io/smm-bench/cascadelake/
https://haampie.github.io/smm-bench/skylake-avx512/ https://haampie.github.io/smm-bench/skylake-avx512/
LoopVectorization.jl performs best with AVX512, but not badly without it:
https://haampie.github.io/smm-bench/znver2/ https://haampie.github.io/smm-bench/znver2/
Versus Eigen, LoopVectorization.jl does better at small sizes (less than a couple hundred, up until packing matters) on my computer.