Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
salykova
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
Matrix Core Programming on AMD CDNA4 Architecture
(rocm.blogs.amd.com)
3 points
by
salykova
10mo ago
|
0 comments
2.
▲
Matrix Core Programming on AMD GPUs
(rocm.blogs.amd.com)
2 points
by
salykova
10mo ago
|
0 comments
3.
▲
Matrix Core Programming on AMD CDNA Architecture
(rocm.blogs.amd.com)
61 points
by
salykova
10mo ago
|
22 comments
4.
▲
Matrix Core Programming on AMD GPUs
(rocm.blogs.amd.com)
1 points
by
salykova
10mo ago
|
0 comments
5.
▲
Matrix Core Programming on AMD CDNA3 and CDNA4 Architecture
(rocm.blogs.amd.com)
1 points
by
salykova
10mo ago
|
0 comments
6.
▲
Matrix Core Programming on AMD CDNA3 and CDNA4 Architecture
(rocm.blogs.amd.com)
2 points
by
salykova
11mo ago
|
0 comments
7.
▲
by
salykova
2y ago
excalidraw <3
8.
▲
by
salykova
2y ago
as we discussed earlier, the code really needs Clang to attain high performance
9.
▲
by
salykova
2y ago
"off-topic" channel
10.
▲
by
salykova
2y ago
We were actively chatting with Justine yesterday, seems like the implementation is at least 2x faster than tinyBLAS on her workstation. The whole discussion is in Mozilla AI discord: https://discord.com/invite/NSnjHmT5x
11.
▲
by
salykova
2y ago
Hi! I'm the author of the article. It's my really first time optimizing C code and using intrinsics, so I'm definitely not an expert in this area, but Im willing to learn more! Many thanks for your feedback; I truly appreciat
12.
▲
Beating NumPy's matrix multiplication in 150 lines of C code
(salykova.github.io)
5 points
by
salykova
2y ago
|
0 comments
13.
▲
Beating NumPy's matrix multiplication in 150 lines of C code
(salykova.github.io)
2 points
by
salykova
2y ago
|
0 comments