3 ms·
Is the author a genius or has AMD questionable software?
by randomNumber7 2y ago
Is the author a genius or has AMD questionable software?
- imtringued 2y agoConsidering the biggest difference between the kernels is the lack of dual issue instructions (an AMD specific innovation). I'd bet on the latter.
- kimixa 2y agoMany of the optimizations here rely heavily on the size of matrix and it's relationship to hardware specific details, like LDS size, how they're banked and register count. It's probably not surprising that you can grind a decent improvement over a general solution, and many of the improvements shown here will need to be re-balanced, or even simply not work, for kernels working on different matrix layouts. Similarly for trying to work on different hardware - even in the same architecture and generation these sort of details are often changing. And all that required going down to the ISA level, which is a lot less easy (certainly less documented) for Nvidia - for example the "inspiration" post linked [0] on CUDA didn't beat cuBLAS also didn't try modifying the SASS directly, so there might be similar level gains unrealized there. [0] https://siboehm.com/articles/22/CUDA-MMM https://siboehm.com/articles/22/CUDA-MMM
- almostgotcaught 2y ago> like LDS size, how they're banked and register count. but you're acting like they pick these numbers using a random number generator for each generation when it's just reasonable/rational stuff like "here's 2x more LDS or more registers for free because the new process node is 2x smaller". like you must realize that they're not throwing everything away and starting completely from scratch for every new gen right? incidentally, while LDS will grow and # of registers will grow, there's absolutely no way they'd change the banking - e.g., CUDA hasn't changed it since 2.0.
- kimixa 2y agoNo, but it's not obviously clear that other sized kernels will hit the same bottlenecks seen in the post. It's not really shown one way or the other - is it that the rocm kernels are just inefficient, or just the author identified one that wasn't particularly well optimized? And do these opportunities for improvement really mean that the software is "Questionable", or just that you cannot really do an equivalent comparison at the level of ISA on other vendor's software stacks? I'm not trying to minimize the work here, it's interesting and a good example of the sort of lengths you can go to in order to squeeze that last little bit of performance out (and again, showing the advantages of public ISA documentation and support for users working at that level), I just took issue to the parent comment seeming to use this work as evidence of a poor baseline.
- roenxi 2y agoROCm multiplies in 4.5ms and the author multiplies in 2.8ms. The naive algorithm is 136ms. I don't think anyone at AMD is losing sleep over this; for a general purpose library this isn't horrible performance. It could be better, hand optimising to specific conditions often is. But as this blog post shows, optimising kernels is the sort of thing that people can do for fun and post blogs about if they care. They don't need AMD to be involved. The problem with ROCm isn't that it only half-utilises the hardware, the problem was that someone trying to write this blog post in 2020 would have had (or at least the probability was rather high) a heading somewhere around implementing Kernel 0 talking about how the software crashed or the kernel panicked when they tried to run the benchmarks. That was what happened to me when I tried a conceptually similar exercise. I was wandering around HN posting comments about how there were no articles like this one to be found for AMD hardware and musing whether it was technically possible to do. This makes me wish I'd bought an RDNA3 card instead of a Nvidia one for my last purchase. Not that I really regret the choice, AMD are going to have to show that they're interested in supporting consumer cards for a little longer to get me to trust them again although they're on the right path.
- latchkey 2y agoFollow Anush on Twitter and give him feedback. He's actively listening. https://x.com/AnushElangovan https://x.com/AnushElangovan
- saagarjha 2y agoAMD isn’t losing sleep over the fact that J. Random Blogger is beating their GEMM by 60% on 4096x4096? What universe are you living in? This company is fighting for their life against CUDA and you’re telling me their software stack being so bad it can’t use a third of the hardware on the the first and literally only thing people want it to do is somehow not a problem?
- roenxi 2y agoThe point of a platform is for software engineers to provide key functionality independently. Your issue here is you don't understand why CUDA has been so dominant over the last decade - a ~50% software performance gap isn't that material when hardware capacity doubles every generation. If we've reached the point where J. Random Blogger can solve their own problems then the CUDA moat has quite possibly been broken. If AMD was only 1 hardware generation behind Nvidia they'd be pretty competitive. People are happy using CPUs with a gap of several generations from the cutting edge. And it isn't even that bad because anyone who particularly cares can optimise their software and avoid using rocBLAS.
- latchkey 2y agoHe used to work for AMD. https://www.linkedin.com/in/sebastienvince/ https://www.linkedin.com/in/sebastienvince/