Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
vishal-padia
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
Show HN: Where a hand-written GPU kernel beats the library, and where it can't
(kyrieblunders.bearblog.dev)
1 points
by
vishal-padia
2mo ago
|
0 comments
2.
▲
I made a kernel 2.2x faster. It made my training loop 3x slower
(kyrieblunders.bearblog.dev)
24 points
by
vishal-padia
4mo ago
|
3 comments
3.
▲
by
vishal-padia
4mo ago
Quick context on what's in the post: 1. From scratch Dr. GRPO implementation in ~300 lines of PyTorch (Qwen2.5-0.5B on GSM8K, A10G). 2. Profiling deep dive on the training loop. Generate is 90% of step time. Pre-allocating the KV cache
4.
▲
Wrote FA2 using CuteDSL, what worked and what didn't work
(kyrieblunders.bearblog.dev)
2 points
by
vishal-padia
5mo ago
|
0 comments
5.
▲
by
vishal-padia
1y ago
working on local first, bloat free MLOps tool https://github.com/Vishal-Padia/tracely it's still wip
6.
▲
by
vishal-padia
2y ago
yes this!!! Whenever I write a prompt, I tend to divide it into smaller prompts, and in this process, my brain thinks of multiple ways to solve the problem. So yes, it's not limiting my thought process. I didn't notice this thing
7.
▲
by
vishal-padia
2y ago
I don't agree, in the initial stages solving problems without LLMs will give a good enough knowledge about the intricacies involved and it helps develop a structured approach while solving a problem!
8.
▲
by
vishal-padia
2y ago
In the article, you mentioned that you've been writing code for 36 years, so don't you feel IDEs like cursor make you feel less competent? Meaning I loved the process of scratching my head over a problem and then coming to a solut