4 ms·
An efficient xgemm kernel is probably faster than the code that taco generates if you don't need to permute your tensor. A contraction like B(i,k) = T(i,j,k) *
by lsorber 9y ago
An efficient xgemm kernel is probably faster than the code that taco generates if you don't need to permute your tensor. A contraction like B(i,k) = T(i,j,k) * A(j) would require a permutation of T before you could run the matrix multiplication though, while taco can just keep the data in-place.