4 ms·
https://www.cs.utexas.edu/~jianyu/papers/sc16.pdf https://www.cs.utexas.edu/~jianyu/papers/sc16.pdf (paper) https://www.cs.utexas.edu/~jianyu/presentations/stra
by jedbrown 8y ago
https://www.cs.utexas.edu/~jianyu/papers/sc16.pdf https://www.cs.utexas.edu/~jianyu/papers/sc16.pdf (paper)
https://www.cs.utexas.edu/~jianyu/presentations/strassen_sc16.pdf https://www.cs.utexas.edu/~jianyu/presentations/strassen_sc1... (slides)
- joshuamorton 8y agoIf I'm reading page 65 of that presentation correctly, naive Strassen (what this person did) should be approximately 75% as fast as MKL on an 8 core machine for a 4Kx4K matrix, and even the improved algorithm outlined in the paper is only approximately equivalent, and I wouldn't call it "Strassen's". I stand by what I said (although that paper was a cool read!)
- jedbrown 8y agoSo ABC Strassen (one of the variants in the paper) is crossing over at 4000x4000 on 10 cores and ahead on fewer cores. I shared this as the current state of performance engineering for Strassen-like algorithms. It offers modest benefits in more practical regimes than the conventional wisdom. I agree that many applications of dense matrices do not benefit.