4 ms·
I'm waiting on a plane, so it's a bit messy, and I haven't tested the float version yet but: https://github.com/RDeckers/c_doodles/blob/master/Linalg/src/linalg
by Coding_Cat 11y ago
I'm waiting on a plane, so it's a bit messy, and I haven't tested the float version yet but: https://github.com/RDeckers/c_doodles/blob/master/Linalg/src/linalg.c https://github.com/RDeckers/c_doodles/blob/master/Linalg/src...
Compile with gcc -O3.
MxM_4x4 is my function, say we have AB=C then it loads B transposed into registers, then for each row it uses 4 multiplications to calculate to row(A)column(B) products, 2 HADD for summing half of the generated products, then 2 permutations to line up the vectors, and one final sum to compute a row of C.
MxM_4x4_2 is a direct port to doubles instead of floats of intel's example on the provided link. When compiled with -O3 my compiler produces the same code in terms of assembly instructions as Intel's example explictlly writes.
(Note the name of the repo, I know it's not pretty ;) )