Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
skidrow
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
61.
▲
Outperforming cuBLAS on H100: A Worklog
(cudaforfun.substack.com)
3 points
by
skidrow
2y ago
|
0 comments
62.
▲
Optimizing Matrix Multiplication on RDNA3
(seb-v.github.io)
118 points
by
skidrow
2y ago
|
26 comments
63.
▲
Outperforming cuBLAS on H100: A Worklog
(cudaforfun.substack.com)
1 points
by
skidrow
2y ago
|
0 comments
64.
▲
Mastering LLM Techniques: Inference Optimization
(developer.nvidia.com)
2 points
by
skidrow
2y ago
|
0 comments
65.
▲
Optimizing Matrix Multiplication on RDNA3
(seb-v.github.io)
2 points
by
skidrow
2y ago
|
0 comments
66.
▲
Outperforming cuBLAS on H100: A Worklog
(cudaforfun.substack.com)
4 points
by
skidrow
2y ago
|
0 comments
67.
▲
Understanding Latency Hiding on GPUs [pdf]
(www2.eecs.berkeley.edu)
2 points
by
skidrow
2y ago
|
0 comments
68.
▲
AMD Radeon RX 9070 Series Linux GPU Compute Performance
(phoronix.com)
2 points
by
skidrow
2y ago
|
0 comments
69.
▲
Outperforming cuBLAS on H100: A Worklog
(cudaforfun.substack.com)
3 points
by
skidrow
2y ago
|
0 comments
70.
▲
GPU Gems
(developer.nvidia.com)
2 points
by
skidrow
2y ago
|
1 comments
71.
▲
Understanding Latency Hiding on GPUs [pdf]
(www2.eecs.berkeley.edu)
2 points
by
skidrow
2y ago
|
0 comments
72.
▲
A guide to LLM inference and performance
(baseten.co)
1 points
by
skidrow
2y ago
|
0 comments
73.
▲
Mastering LLM Techniques: Inference Optimization
(developer.nvidia.com)
2 points
by
skidrow
2y ago
|
0 comments
74.
▲
GPT from Scratch with MLX
(rayfernando.ai)
1 points
by
skidrow
2y ago
|
0 comments
75.
▲
Mastering LLM Techniques: Inference Optimization
(developer.nvidia.com)
3 points
by
skidrow
2y ago
|
0 comments
76.
▲
Beating OpenBLAS in FP32 Matrix Multiplication
(salykova.github.io)
4 points
by
skidrow
2y ago
|
1 comments
77.
▲
Beating OpenBLAS in FP32 Matrix Multiplication
(salykova.github.io)
1 points
by
skidrow
2y ago
|
0 comments
78.
▲
GPT from Scratch with MLX
(rayfernando.ai)
1 points
by
skidrow
2y ago
|
0 comments
79.
▲
Mastering LLM Techniques: Inference Optimization
(developer.nvidia.com)
1 points
by
skidrow
2y ago
|
0 comments
80.
▲
Mastering LLM Techniques: Inference Optimization
(developer.nvidia.com)
1 points
by
skidrow
2y ago
|
0 comments
81.
▲
Beating OpenBLAS in Matrix Multiplication
(salykova.github.io)
1 points
by
skidrow
2y ago
|
0 comments
82.
▲
GPT from Scratch with MLX
(rayfernando.ai)
3 points
by
skidrow
2y ago
|
0 comments
83.
▲
An overview of gradient descent optimization algorithms (2016)
(ruder.io)
135 points
by
skidrow
2y ago
|
26 comments
84.
▲
An overview of gradient descent optimization algorithms
(ruder.io)
1 points
by
skidrow
2y ago
|
0 comments
85.
▲
An overview of gradient descent optimization algorithms
(ruder.io)
2 points
by
skidrow
2y ago
|
0 comments
86.
▲
Pipeline-Parallelism: Distributed Training via Model Partitioning
(siboehm.com)
1 points
by
skidrow
2y ago
|
0 comments
87.
▲
Beating OpenBLAS in FP32 Matrix Multiplication
(salykova.github.io)
2 points
by
skidrow
2y ago
|
0 comments
88.
▲
Beating OpenBLAS in FP32 Matrix Multiplication
(salykova.github.io)
2 points
by
skidrow
2y ago
|
0 comments
89.
▲
Beating cuBLAS in Single-Precision General Matrix Multiplication
(salykova.github.io)
3 points
by
skidrow
2y ago
|
0 comments
90.
▲
Beating OpenBLAS in FP32 Matrix Multiplication
(salykova.github.io)
4 points
by
skidrow
2y ago
|
0 comments
More ›