7 ms·
My experience with it was: with CSR format, SpMV got about 40x speedup compared to a single CPU core on a RTX 3090, with full core utilization. Rough calculatio
by syockit 3y ago
My experience with it was: with CSR format, SpMV got about 40x speedup compared to a single CPU core on a RTX 3090, with full core utilization. Rough calculation tells me it's equivalent to 80 GFlops, which is far from the full potential of the card. I thought there was someting wrong with cuSPARSE so I went ahead and wrote a SpMV kernel, based on the code in NVIDIA's past reports. It actually performed slightly better. Fair enough, cuSPARSE has to support all kind of workloads so it might be just that in my case, the handwritten kernel worked better. But still, it would never get to the regime of hundreds of GFlops. After reading various literatures on GPU SpMV implementations getting to the same conclusion, I had to concede that I can't squeeze more GFlops out of it.
- KeplerBoy 3y agoHave you checked the roofline performance model as reported by Nsight Compute? A lot of workloads can never reach the hundreds of GFlops because they are bandwidth limited.
- syockit 3y agoI couldn't check it using Nsight Compute. It didn't find the kernel even though I ran it as root. A40 is supposed to be Ampere architecture, with GA102 chip so it should be supported. Maybe the toolkit I'm using is too old (CUDA 11.3). Everything is on the device during the tight loop, and there is minimal transfer to host, so I don't think it's bandwidth limited in that sense. It could be bandwidth limited on the cache level but I couldn't tell. I wish I could get Compute running.