Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
skidrow
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
5 ms
·
31.
▲
FP8 GEMM Optimization on AMD CDNA4 Architecture
(rocm.blogs.amd.com)
3 points
by
skidrow
4mo ago
|
0 comments
32.
▲
Deep Dive into 4-Wave Interleave FP8 GEMM
(rocm.blogs.amd.com)
3 points
by
skidrow
4mo ago
|
0 comments
33.
▲
Occupancy Math on the AMD MI355X: A From-First-Principles Guide
(indianspeedster.github.io)
3 points
by
skidrow
4mo ago
|
0 comments
34.
▲
Creating custom kernels for the AMD MI300
(huggingface.co)
2 points
by
skidrow
1y ago
|
0 comments
35.
▲
Implementing a Fast Tensor Core Matmul on the Ada Architecture
(spatters.ca)
4 points
by
skidrow
1y ago
|
0 comments
36.
▲
Matrix Core Programming on AMD GPUs
(salykova.github.io)
116 points
by
skidrow
1y ago
|
5 comments
37.
▲
Implementing a Fast Tensor Core Matmul on the Ada Architecture
(spatters.ca)
3 points
by
skidrow
1y ago
|
0 comments
38.
▲
Matrix Core Programming on AMD GPUs
(salykova.github.io)
2 points
by
skidrow
1y ago
|
0 comments
39.
▲
Creating custom kernels for the AMD MI300
(huggingface.co)
1 points
by
skidrow
1y ago
|
0 comments
40.
▲
Implementing a Fast Tensor Core Matmul on the Ada Architecture
(spatters.ca)
2 points
by
skidrow
1y ago
|
0 comments
41.
▲
Matrix Core Programming on AMD CDNA3 and CDNA4 Architecture
(salykova.github.io)
24 points
by
skidrow
1y ago
|
3 comments
42.
▲
Creating custom kernels for the AMD MI300
(huggingface.co)
2 points
by
skidrow
1y ago
|
0 comments
43.
▲
Implementing a Fast Tensor Core Matmul on the Ada Architecture
(spatters.ca)
2 points
by
skidrow
1y ago
|
0 comments
44.
▲
Advanced Matrix Multiplication Optimization on Multi-Core Processors (2024)
(salykova.github.io)
85 points
by
skidrow
1y ago
|
3 comments
45.
▲
Creating custom kernels for the AMD MI300
(huggingface.co)
2 points
by
skidrow
1y ago
|
0 comments
46.
▲
Introduction to Matrix Core Programming on AMD CDNA3 and CDNA4 Architecture
(salykova.github.io)
2 points
by
skidrow
1y ago
|
0 comments
47.
▲
Creating custom kernels for the AMD MI300
(huggingface.co)
2 points
by
skidrow
1y ago
|
0 comments
48.
▲
Implementing a Fast Tensor Core Matmul on the Ada Architecture
(spatters.ca)
2 points
by
skidrow
1y ago
|
0 comments
49.
▲
Creating custom kernels for the AMD MI300
(huggingface.co)
1 points
by
skidrow
1y ago
|
0 comments
50.
▲
Implementing a Fast Tensor Core Matmul on the Ada Architecture
(spatters.ca)
4 points
by
skidrow
1y ago
|
0 comments
51.
▲
Implementing a Fast Tensor Core Matmul on the Ada Architecture
(spatters.ca)
2 points
by
skidrow
1y ago
|
1 comments
52.
▲
Compiler Explorer: An Essential Kernel Playground for CUDA Developers
(developer.nvidia.com)
2 points
by
skidrow
1y ago
|
0 comments
53.
▲
Creating custom kernels for the AMD MI300
(huggingface.co)
1 points
by
skidrow
1y ago
|
0 comments
54.
▲
DeepSeek-R1 and FP8 Mixed-Precision Training
(research.colfax-intl.com)
2 points
by
skidrow
1y ago
|
0 comments
55.
▲
How to Write a Fast Matrix Multiplication from Scratch with Tensor Cores (2024)
(alexarmbr.github.io)
147 points
by
skidrow
1y ago
|
17 comments
56.
▲
DeepSeek-R1 and FP8 Mixed-Precision Training
(research.colfax-intl.com)
2 points
by
skidrow
1y ago
|
0 comments
57.
▲
Implementing a Fast Tensor Core Matmul on the Ada Architecture
(spatters.ca)
1 points
by
skidrow
1y ago
|
0 comments
58.
▲
How to Write a Fast Matrix Multiplication from Scratch with Tensor Cores
(alexarmbr.github.io)
2 points
by
skidrow
1y ago
|
0 comments
59.
▲
Understanding Peak, Max-Achievable and Delivered FLOPs
(rocm.blogs.amd.com)
1 points
by
skidrow
2y ago
|
0 comments
60.
▲
DeepSeek-R1 and FP8 Mixed-Precision Training
(research.colfax-intl.com)
1 points
by
skidrow
2y ago
|
0 comments
More ›