Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
jaberjaber23
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
6 ms
·
1.
▲
by
jaberjaber23
8mo ago
interesting mix of k and joy. the queue manipulation primitives like -> and => have no equivalent in joy, lets you do things like call/cc in a few lines
2.
▲
by
jaberjaber23
8mo ago
seconding the jamis buck book, its one of the few programming books i actually finished. the way he explains each algorithm with visualizations makes it stick
3.
▲
by
jaberjaber23
8mo ago
had the same thing, electrician just used whatever was on the truck. swapped the faceplates and got instant gigabit
4.
▲
by
jaberjaber23
9mo ago
Six months ago, if you asked me whether an LLM could write a CUDA kernel that actually beats PyTorch's compiler, I would have said no. The optimization space is too complex. Too many hardware details. Too easy to write something that c
5.
▲
by
jaberjaber23
9mo ago
amazing paper
6.
▲
by
jaberjaber23
11mo ago
Wow man, will give it a try
7.
▲
by
jaberjaber23
11mo ago
Impressive work, will give it a try!!
8.
▲
by
jaberjaber23
11mo ago
I was using 4 sometimes 6 different tools just to write CUDA. vs code for coding, nsight for profiling, many custom tools for benchmarking and debugging, plus pen to calc the performance "I was cooked" So I built code editor for C
9.
▲
by
jaberjaber23
1y ago
science repeats itself
10.
▲
by
jaberjaber23
1y ago
Nanopore’s getting closer
11.
▲
by
jaberjaber23
1y ago
amazing!!
12.
▲
by
jaberjaber23
1y ago
most people don’t procrastinate because they’re lazy, they procrastinate because their brain rejects meaningless work
13.
▲
by
jaberjaber23
1y ago
llms don’t actually freak out over seahorses. it’s just a mismatch. the model thinks “seahorse emoji” is real, but the output system doesn’t have a token for it. it tries to show what it means, realizes it can’t, and spirals trying to fix i
14.
▲
by
jaberjaber23
1y ago
testing cuda kernels on different gpus costs $7k/month in cloud rentals so built an emulator instead you give it a kernel, it predicts execution time on any gpu without running it. h100, a100, v100, whatever. how: scraped specs for 50+
15.
▲
by
jaberjaber23
1y ago
Testing CUDA kernels on 15 GPUs costs thousands every month we couldn’t afford that, so we built an emulator that predicts how your kernel runs on any GPU like H100, A100, 4090, or V100 without running a single line it’s not a guess, it giv
16.
▲
by
jaberjaber23
1y ago
Testing CUDA kernels on 15 different GPUs costs $3,000/month. We couldn't afford that!! So we built an emulator. give it your kernel code and it tells you exactly how it runs on any GPU. H100, A100, RTX 4090, V100, whatever "
17.
▲
by
jaberjaber23
1y ago
exactly
18.
▲
by
jaberjaber23
1y ago
Arabic AI training data is terrible. Everything is either tiny or garbage quality. So I built my own. 744K articles, 244 million words. Spent months cleaning it properly. It's 8.7GB of good Arabic text covering everything. Made it comp
19.
▲
by
jaberjaber23
1y ago
I built one of the largest Arabic LLM datasets: - 743K articles - 244M words - 1.5M unique words - Cleaned, deduplicated, JSONL format, LLM-ready It’s ready for GPT, BERT, LLaMA, or any Arabic NLP research. For too long, Arabic AI lagged be
20.
▲
by
jaberjaber23
1y ago
AI hits a scaling wall More GPUs → smaller gains That’s the law. The only move left: shift the intercept. Make the same FLOPs buy lower loss. I broke it down: - Why scaling stalls - What actually shifts the curve - A one-week playbook to ru
21.
▲
by
jaberjaber23
1y ago
I’ve spent years manually tuning CUDA kernels, and it’s exhausting. Recently, I experimented with real-time profiling to see bottlenecks instantly. The results were surprising—small tweaks yielded huge performance gains. This got me thinkin
22.
▲
by
jaberjaber23
1y ago
Just open-sourced CUDA CLI :) It automatically optimizes CUDA kernels. Depending on the kernel, you can get 10–50x speedups :D How it works: - Finds bottlenecks in your kernel - Generates a few optimized variants - Benchmarks them - Keeps t
23.
▲
Open Source AI CUDA Kernel Optimizer
(github.com)
2 points
by
jaberjaber23
1y ago
|
1 comments
24.
▲
by
jaberjaber23
1y ago
RightNow CLI is now open source. It automatically optimizes CUDA kernels using AI. Depending on the kernel, it can achieve 10 to 50x speedups by removing bottlenecks and tuning for your GPU. How it works: - Analyze kernel patterns - Generat
25.
▲
by
jaberjaber23
1y ago
Most developers guess why their CUDA code is slow. I wanted answers, so I built RightNow AI. It’s a GPU-first editor that: Profiles your kernels live without leaving the editor Lets AI rewrite and tune them for your exact GPU Autocompletes
26.
▲
by
jaberjaber23
1y ago
Most CUDA developers waste hours trying to speed up kernels because they can’t see the full architecture and behavior of their GPU. So, I built RightNow AI solo while still at university. It’s a CUDA code editor that profiles kernels live i
27.
▲
by
jaberjaber23
1y ago
what's up guys, take it easy. Just to clarify: I didn’t add any reviews myself. I’m building this SOLO and barely have time to finish the product, let alone fake comments. I wasn’t even aware of the Product Hunt stuff until people here
28.
▲
by
jaberjaber23
1y ago
I really appreciate that!! thanks:D
29.
▲
by
jaberjaber23
1y ago
Do you have a place where we can chat? Linkedin,....
30.
▲
by
jaberjaber23
1y ago
absolutely. it really depends on the kernel type, target architecture, and what you're optimizing for. the 2x-4x isn’t the limit, it's just what users often see out of the box. we do real-time profiling on actual GPUs, so you get
More ›