Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
ap4
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
by
ap4
2y ago
I have deleted it. The function verify_matrix() from the original SGEMM_CUDA repository did not check for NANs, and the kernel was returning NANs. I have no way to delete this submission.
2.
▲
by
ap4
2y ago
You are right. The function verify_matrix() from the original SGEMM_CUDA repository did not check for NANs. I deleted the repository. It was the 13th CUDA kernel I wrote in my life, and the whole endeavor teached me a lot. I appreciate the
3.
▲
by
ap4
2y ago
I want to have 512 threads per block, each thread calculating simultaneously 128 values. That's 65536 values per block. I can't accumulate each of these values in registers, because the GPU has the limit of max 65536 registers per
4.
▲
by
ap4
2y ago
That's the trick. This is intended. The point is that the compiler does not notice oob access in the first stage, but notices it in the later stages, and compiles the code to a correctly working kernel. The result is correct, as checke
5.
▲
by
ap4
2y ago
For all doubters, I open-sourced it: https://github.com/arekpaterek/Faster_SGEMM_CUDA
6.
▲
by
ap4
2y ago
I shared the code: https://github.com/arekpaterek/Faster_SGEMM_CUDA
7.
▲
by
ap4
2y ago
I did more tests on various GPUs. The TFLOPS values for A100 exceed the maximum from the spec of the GPU, so perhaps TFLOPS is calculated in a different way in the spec. size tflops_cublas tflops_my diff gpu 12288² 51.4
8.
▲
by
ap4
2y ago
Tested on more GPUs. The biggest improvement found so far over the standard matrix multiplication from the cuBLAS library is +43% when multiplying matrices of size 12288² on an A100 GPU. size tflops_cublas tflops_my diff gpu
9.
▲
by
ap4
2y ago
Also on other GPUS. So far I have tested on: size tflops_cublas tflops_my diff gpu 4096² 28.7-28.8 32.5 +13% 4070ts 8192² 27.7-28.2 33.5 +19-21% 4070ts 4096² 9.9-10.0 10.1-10.2 +
10.
▲
by
ap4
2y ago
Related: I created a CUDA kernel typically much faster than kernels from cuBLAS when multiplying large square float32 matrices. Tested mostly on a 4090 GPU so far. Source code: https://github.com/arekpaterek/Faster_SGEM
11.
▲
Show HN: FP32 matmul of large matrices up to 24% faster than cuBLAS on a 4090
(github.com)
4 points
by
ap4
2y ago
|
4 comments
12.
▲
by
ap4
2y ago
Believe it or not. On a 4090 gpu, average of 20 runs of SGEMM_CUDA: size tflops_cublas tflops_my diff 4096² 50.8-50.9 61.8 +21% 8192² 56.3-56.4 67.1 +19% 16384² 53.6 66.7 +24% I g
13.
▲
by
ap4
2y ago
It is apples-to-apples. The code from the article runs a chosen kernel 50 times in a loop. In the case of CuBLAS it runs the function cublasGemmEx() 50 times.
14.
▲
by
ap4
2y ago
I swapped one of the kernels in the code from the article to my kernel, and left only the multiplication of matrices of size 4096². On average over 20 runs: CuBLAS (./sgemm 0) has 50.9 TFLOPS. My kernel has 61.8 TFLOPS, so it's ac
15.
▲
by
ap4
2y ago
Beating CuBLAS is easy. I wrote a 30-line kernel that multiplies two 4096² matrices 14% faster than CuBLAS on a 4090. The question is how to earn money on that.