Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
tspeterkim
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
I came back to school to study hardware after 5 years of doing ML
(tspeterkim.github.io)
2 points
by
tspeterkim
2y ago
|
0 comments
2.
▲
I came back to school to study hardware after 5 years of doing ML
(tspeterkim.github.io)
1 points
by
tspeterkim
2y ago
|
0 comments
3.
▲
How to speed up Flux inference by 40% in one line
(twitter.com)
1 points
by
tspeterkim
2y ago
|
0 comments
4.
▲
A Minimal KV Cache Manager for Paged Attention in ~100 Lines of Python
(github.com)
2 points
by
tspeterkim
2y ago
|
0 comments
5.
▲
Show HN: Minimal Paged Attention
(github.com)
3 points
by
tspeterkim
2y ago
|
0 comments
6.
▲
Show HN: DIY Instagram Automation for My Influencer Wife
(github.com)
3 points
by
tspeterkim
2y ago
|
3 comments
7.
▲
Show HN: Mixed Precision Training from Scratch
(github.com)
1 points
by
tspeterkim
2y ago
|
0 comments
8.
▲
Mixed Precision Training from Scratch
(tspeterkim.github.io)
2 points
by
tspeterkim
2y ago
|
0 comments
9.
▲
Profiling Remote GPUs with Nsight Compute
(tspeterkim.github.io)
1 points
by
tspeterkim
2y ago
|
0 comments
10.
▲
by
tspeterkim
2y ago
I tried this out today. While it works (no longer a pre-split step required), it makes the CUDA kernel run ridiculously slow. I believe it's because of the while loop: while (i < end_byte) { Comparing it to my original soluti
11.
▲
by
tspeterkim
2y ago
Please do. My hope with this blog was to rile up other CUDA enthusiasts. Making them wanna bring in bigger and better hardware that I don't have access to. + Someone in the CUDA MODE community got 6 seconds on a 4090.
12.
▲
by
tspeterkim
2y ago
This is the possible optimization that I mention at the end of the blog - using a private map for each thread block. The catch is that this map must fit in shared memory, which is pretty limited on all current hardware: ~100KB. I originally
13.
▲
by
tspeterkim
2y ago
I misspoke. (got confused with the key limit in my link above) Atomics work up to 128-bits ( https://docs.nvidia.com/cuda/cuda-c-programming-guide/#atomi... ). Regardless, it's still less than 100 bytes, which
14.
▲
by
tspeterkim
2y ago
for (int i = 0; buffer[offset+i] != '\n'; i++) { This would only process the current line, though. Here, each thread processes ~split_size bytes (multiple lines). Even if were to read multiple lines, how would a thread kno
15.
▲
by
tspeterkim
2y ago
By "launcher", do you mean the CUDA kernel? How can it avoid accessing the input data since it needs access to the characters based on the offsets? I also already pass these offsets to the threads as `Part* parts`. I also probably
16.
▲
by
tspeterkim
2y ago
The PMPP book is great. I reread the histogram chapter after finishing the blog, and realized I could use privatization. You got me! By coarsening, do you mean making the threads handle more file parts, and reducing the number of private co
17.
▲
by
tspeterkim
2y ago
So performance would increase since hashing is faster than binary-searching. However, the problem of collisions across threads and dealing with concurrent map key insertions still remains. e.g. when two different cities produce the same has
18.
▲
by
tspeterkim
2y ago
Agreed. I tried reducing across cities first. The problem was that the work of gathering all the temperatures for each city (before I could launch the reduction CUDA kernels) required a full parsing through the input data. My final solution
19.
▲
Show HN: One Billion Rows in CUDA
(github.com)
3 points
by
tspeterkim
2y ago
|
0 comments
20.
▲
The One Billion Row Challenge in CUDA
(tspeterkim.github.io)
241 points
by
tspeterkim
2y ago
|
74 comments
21.
▲
by
tspeterkim
3y ago
Thanks Daniel. The main blocker is me not able to fully grasp the backward pass. (trying to understand Appendix B.2 in the original paper) I need to get more comfortable with matrix derivatives before I can confidently reimplement it in the
22.
▲
Show HN: Flash Attention in ~100 lines of CUDA
(github.com)
230 points
by
tspeterkim
3y ago
|
39 comments