5 ms·
Related: I created a CUDA kernel typically much faster than kernels from cuBLAS when multiplying large square float32 matrices. Tested mostly on a 4090 GPU so f
by ap4 2y ago
Related: I created a CUDA kernel typically much faster than kernels from cuBLAS when multiplying large square float32 matrices. Tested mostly on a 4090 GPU so far.
Source code: https://github.com/arekpaterek/Faster_SGEMM_CUDA https://github.com/arekpaterek/Faster_SGEMM_CUDA
size tflops_cublas tflops_my diff gpu
4096² 50.8-50.9 61.8 +21% 4090
6144² 55.3 59.8 +8% 4090
8192² 56.3-56.5 67.1 +19% 4090
12288² 53.7 66.7 +24% 4090
16384² 53.6 66.7 +24% 4090
4096² 28.7-28.8 32.5 +13% 4070ts
4096² 3.8-4.3 6.7 +56-76% T4
- ap4 2y agoTested on more GPUs. The biggest improvement found so far over the standard matrix multiplication from the cuBLAS library is +43% when multiplying matrices of size 12288² on an A100 GPU. size tflops_cublas tflops_my diff gpu 12288² 51.4 56.3 +9% h100 8192² 50.5 56.1 +11% h100 4096² 43.8 53.9 +23% h100 12288² 18.9 27.0 +43% a100 8192² 19.0 26.3 +38% a100 4096² 17.5 19.8 +13% a100 12288² 28.8 34.5 +20% 3090ti Edit: The values for A100 are fishy. They should not exceed 19.5 TFLOPS, which is the maximum from the spec. I have to take a closer look into that.
- ladberg 2y agoYour `sum` array is only 64 elements but you're indexing with indices out of bounds, which is UB and the compiler knows it at compile time so it's skipping a bunch of work. E.g. consider the line: sum[y2*8 + x2] += ... In the final loop iteration when y2=15 and x2=7, the index is 127.
- ap4 2y agoThat's the trick. This is intended. The point is that the compiler does not notice oob access in the first stage, but notices it in the later stages, and compiles the code to a correctly working kernel. The result is correct, as checked by the function verify_matrix().
- ladberg 2y agoSorry I'm a bit lost here, could you explain the reasoning behind this and why it works?
- ap4 2y agoI want to have 512 threads per block, each thread calculating simultaneously 128 values. That's 65536 values per block. I can't accumulate each of these values in registers, because the GPU has the limit of max 65536 registers per block, and some additional registers are needed in the kernel. But if I find a way to trick the first stages of the compiler that it has sufficient amount of free registers, then sometimes, like in the case of this kernel, the later stages of the compiler are sufficiently smart to give me what I want: 512 threads per block, each calculating 128 values.
- deleted 2y ago[deleted]
- ladberg 2y agoI hate to say it but that simply doesn't work: you can't write out of bounds to trick the compiler, it'll just ignore your out of bounds work. You can look at the generated sass on godbolt: https://cuda.godbolt.org/z/19excTxM3 https://cuda.godbolt.org/z/19excTxM3 Note that there are 1024 FFMA instructions in the loop but you would expect 16*8*BK = 2048. This would suggest half the operations are skipped, which lines up with the half of writes that are out of bounds being omitted. After the compute loop when you're calculating the final result and storing it, you can see that the FFMAs referencing out of bounds indices write QNAN instead of any real results. Is it possible that the NANs are what are messing with your tests? Those are notoriously hard to deal with correctly, but you should assert that the result doesn't have any NANs whatsoever.
- ap4 2y agoYou are right. The function verify_matrix() from the original SGEMM_CUDA repository did not check for NANs. I deleted the repository. It was the 13th CUDA kernel I wrote in my life, and the whole endeavor teached me a lot. I appreciate the feedback.