3 ms·
It is apples-to-apples. The code from the article runs a chosen kernel 50 times in a loop. In the case of CuBLAS it runs the function cublasGemmEx() 50 times.
by ap4 2y ago
It is apples-to-apples. The code from the article runs a chosen kernel 50 times in a loop. In the case of CuBLAS it runs the function cublasGemmEx() 50 times.
- david-gpu 2y agoDoes your implementation support the complete feature set of cublasGemmEx()? E.g. varying matrix sizes, etc. Does your implementation beat cublas on a wide range of cases, or only cherry-picked examples? Does the code from the article follow the best coding practices for cublas in the first place? Generalizing from a micro benchmark is typically hubris.