Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
aaa370
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
How to Write a Fast CUDA Matrix Multiplication with Nvidia Tensor Cores
(alexarmbr.github.io)
2 points
by
aaa370
2y ago
|
0 comments
2.
▲
by
aaa370
2y ago
I gotta see it to believe it ;)
3.
▲
by
aaa370
2y ago
Another point to consider here is that this project of writing a cuBLAS level GEMM kernel becomes much more challenging if you are doing it with fp16, and are thus competing with the cuBLAS kernels that use tensor cores. The (theoretical) a