3 ms·
Could you expand on the 24x24 unrolling with AVX? How big of a benefit did you obtain, compared to say using a library such as Armadillo, Fastor, Eigen, etc?
by lldbg 6y ago
Could you expand on the 24x24 unrolling with AVX? How big of a benefit did you obtain, compared to say using a library such as Armadillo, Fastor, Eigen, etc?
- Const-me 6y agoEigen was my baseline. As far as I remember, my best AVX version (which unrolls 24x4 blocks, running 6 iterations of the loop) was about 50% faster. Eigen is designed in the assumption these vectors/matrices are large. 24 is not large, fits in 3 vector registers out of the 16 available. Eliminating data access latency of the accumulators is what caused the speedup.
- lldbg 6y agoIn my experience Eigen is not very good either for small matrices, whilst Fastor performs much better. Interesting
- Const-me 6y agoLibraries don’t cut it, IMO. Sometimes you want your complete input data in registers, 4x4 matrices fit there nicely. Other times you want to stream stuff from memory, the above-mentioned 24x24 multiplication streams the complete matrix, while loading 4-long chunks of the vector. I’ve wrote an introduction article some time ago: http://const.me/articles/simd/simd.pdf http://const.me/articles/simd/simd.pdf Here’s more practical one: https://stackoverflow.blog/2020/07/08/improving-performance-with-simd-intrinsics-in-three-use-cases/ https://stackoverflow.blog/2020/07/08/improving-performance-...