3 ms·
Can you explain why you did the naive algorithm here and not any of the fast matrix multiplication ones that trade multiplications for more additions? Just for
by inglor 2y ago
Can you explain why you did the naive algorithm here and not any of the fast matrix multiplication ones that trade multiplications for more additions? Just for educational purposes or is there a performance benefit in the technique?
- deleted 2y ago[deleted]
- zanussbaum 2y agoat least on my m2, the compiled kernel ends up using fast math anyways so using WGSL's fma didn't change anything about the actual kernel that gets run
- hedgehog 2y agoinglor is probably referring to Strassen or Coppersmith–Winograd.
- zanussbaum 2y agooh in that case it was because i didn't know about them :) something to try next!
- wbl 2y agoLast I checked the extra mems really hurt on a lot of cases especially for the more complex ones, but I'm no expert.
- saagarjha 2y agoBecause those algorithms are generally not worth implementing even though their algorithmic complexity is theoretically lower.