5 ms·
Optimizing Matrix Multiplication on RDNA3
- deleted 2y ago[deleted]
- SavageNoble 2y agoThis is really cool. 60% is no joke and as a 7900XTX owner I would love the performance boost. Well done!
- deleted 2y ago[deleted]
- throawayonthe 2y ago[dead]
- almostgotcaught 2y ago> Furthermore, performing custom ISA optimizations makes these changes RDNA3-specific this is overblown at least wrt forward compatibility - all of the instructions used are in RDNA4 and most of them are even in CDNA3 (CDNA4 isn't public yet?) and the ones that aren't exactly there are only slightly renamed (ds_load -> ds_read). Sure it's annoying but it's not the end of the world to have some `#ifdef`s in your code (that's not very much different from what the compiler itself is going to do anyway).
- imtringued 2y agoYou're making the assumption that every kernel developer has enough AMD GPUs from different eras that they can test their ifdefs on all the possible ISAs.
- randomNumber7 2y agoIs the author a genius or has AMD questionable software?
- imtringued 2y agoConsidering the biggest difference between the kernels is the lack of dual issue instructions (an AMD specific innovation). I'd bet on the latter.
- kimixa 2y agoMany of the optimizations here rely heavily on the size of matrix and it's relationship to hardware specific details, like LDS size, how they're banked and register count. It's probably not surprising that you can grind a decent improvement over a general solution, and many of the improvements shown here will need to be re-balanced, or even simply not work, for kernels working on different matrix layouts. Similarly for trying to work on different hardware - even in the same architecture and generation these sort of details are often changing. And all that required going down to the ISA level, which is a lot less easy (certainly less documented) for Nvidia - for example the "inspiration" post linked [0] on CUDA didn't beat cuBLAS also didn't try modifying the SASS directly, so there might be similar level gains unrealized there. [0] https://siboehm.com/articles/22/CUDA-MMM https://siboehm.com/articles/22/CUDA-MMM
- almostgotcaught 2y ago> like LDS size, how they're banked and register count. but you're acting like they pick these numbers using a random number generator for each generation when it's just reasonable/rational stuff like "here's 2x more LDS or more registers for free because the new process node is 2x smaller". like you must realize that they're not throwing everything away and starting completely from scratch for every new gen right? incidentally, while LDS will grow and # of registers will grow, there's absolutely no way they'd change the banking - e.g., CUDA hasn't changed it since 2.0.
- kimixa 2y ago
- nyanpasu64 2y agoIs it worth implementing sub-cubic matrix multiplication algorithms like Strassen etc. for 4096x4096?
- saagarjha 2y agoI don't think anyone really does this, at least on the GPU.
- spookie 2y agoDependent on your case, but yes, even for smaller matrices.
- 1W6MIC49CYX9GAP 2y agoNo
- tgtweak 2y agoCuda has similar inefficiencies and many use cases can have equal uplifts by going lower level on the code. I think this is what deepseek had done to get their speedups on older hardware. Even way back in the days of GPU crypto mining - custom kernels hand built (mostly just unrolling loops) would yield 20% improvements over just running opencl and letting the drivers compile it down.
- touisteur 2y agoPeople have been trying to bypass CUDA and even PTX for a long time. One long rundown of optimizing gemm on NVIDIA hardware (https://salykova.github.io/sgemm-gpu https://salykova.github.io/sgemm-gpu) mentions 'maxas' (https://github.com/NervanaSystems/maxas/wiki/Introduction https://github.com/NervanaSystems/maxas/wiki/Introduction) - which was really a step forward in this space. I still blame Intel (buying NervanaSystems) for killing it...
- almostgotcaught 2y ago> People have been trying to bypass CUDA and even PTX for a long time i swear it's so funny when people talk about this stuff like it's all weird/surprising. y'all realize that there are hundreds (thousands?) of engineers across FAANG whose full time job is optimizing CUDA/ROCm/whatever code for their team/org/company's specific workloads? like do y'all think that serious shops really just go with whatever the vendor gives you? ie none of this is in the least surprising - it's completely expected that whatever the vendor designs generically for the entire market segment will fail to achieve peak perf for your use case.
- cma 2y ago>it's completely expected that whatever the vendor designs generically for the entire market segment will fail to achieve peak perf for your use case. When Carmack left Meta I believe he claimed they were only getting around 20% utilization on their even then enormous GPU fleet. So I could see them also leaving a lot of perf headroom on the table.
- 2y ago
- delusional 2y agoI find it quite interesting that while vector instructions are present every other sort of "hardware level grouping" (wave, SIMD, thread) is hidden from the programmer. Why would vector instructions be the only thing the programmer ought to care about? I wonder if there's untapped potential in a GPU language which made all of those implicit classes explicit in code, now that we've sort of stabilized on them. It wouldn't allow you to do anything that you can't already do with clever optimizations and a profiler, but it could have the potential to make the optimizations clearer. In general I'm very curious as to why we don't have any new languages that are better aligned with current hardware. For some reason we collectively decided that it was more fun to make everything general, which is especially unfortunate considering the real world got increasingly homogeneous. Compiling to some intermediate language makes no sense when you're only ever going to run on x86 anyway.