4 ms·
Beating CuBLAS is easy. I wrote a 30-line kernel that multiplies two 4096² matrices 14% faster than CuBLAS on a 4090. The question is how to earn money on that.
by ap4 2y ago
Beating CuBLAS is easy. I wrote a 30-line kernel that multiplies two 4096² matrices 14% faster than CuBLAS on a 4090.
The question is how to earn money on that.
- coffeeaddict1 2y agoI'm a little sceptical of your claims. Care to share the kernel you wrote?
- david-gpu 2y agoI would bet actual money that they are not doing an apples-to-apples comparison. I have seen how those high-performance libraries are made and I'm still in awe at the quality and quantity of the staffing involved. Those were the smartest and most knowledgeable engineers I met in my career.
- ap4 2y agoIt is apples-to-apples. The code from the article runs a chosen kernel 50 times in a loop. In the case of CuBLAS it runs the function cublasGemmEx() 50 times.
- david-gpu 2y agoDoes your implementation support the complete feature set of cublasGemmEx()? E.g. varying matrix sizes, etc. Does your implementation beat cublas on a wide range of cases, or only cherry-picked examples? Does the code from the article follow the best coding practices for cublas in the first place? Generalizing from a micro benchmark is typically hubris.
- ap4 2y agoI shared the code: https://github.com/arekpaterek/Faster_SGEMM_CUDA https://github.com/arekpaterek/Faster_SGEMM_CUDA
- ladberg 2y agoIf this were true (and I highly doubt it) it's obvious how to make money from it: collect a 7 figure paycheck from Nvidia, AMD, or any FAANG.
- ap4 2y agoI swapped one of the kernels in the code from the article to my kernel, and left only the multiplication of matrices of size 4096². On average over 20 runs: CuBLAS (./sgemm 0) has 50.9 TFLOPS. My kernel has 61.8 TFLOPS, so it's actually +21% speedup in this benchmark. How do I collect my paycheck?
- JuanJohnJames 2y agoPost the code and your curriculum
- aaa370 2y agoI gotta see it to believe it ;)
- ap4 2y agoBelieve it or not. On a 4090 gpu, average of 20 runs of SGEMM_CUDA: size tflops_cublas tflops_my diff 4096² 50.8-50.9 61.8 +21% 8192² 56.3-56.4 67.1 +19% 16384² 53.6 66.7 +24% I guess the right thing to do now would be to hire a B2B salesman and figure out, which company needs it.
- ap4 2y agoFor all doubters, I open-sourced it: https://github.com/arekpaterek/Faster_SGEMM_CUDA https://github.com/arekpaterek/Faster_SGEMM_CUDA