4 ms·
I couldn't check it using Nsight Compute. It didn't find the kernel even though I ran it as root. A40 is supposed to be Ampere architecture, with GA102 chip so
by syockit 3y ago
I couldn't check it using Nsight Compute. It didn't find the kernel even though I ran it as root. A40 is supposed to be Ampere architecture, with GA102 chip so it should be supported. Maybe the toolkit I'm using is too old (CUDA 11.3).
Everything is on the device during the tight loop, and there is minimal transfer to host, so I don't think it's bandwidth limited in that sense. It could be bandwidth limited on the cache level but I couldn't tell. I wish I could get Compute running.