Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
zanussbaum
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
The Lost Nuance of Grep vs. Semantic Search
(nuss-and-bolts.com)
2 points
by
zanussbaum
11mo ago
|
0 comments
2.
▲
I Like Big Batches and I Cannot Lie: Tricks for Good Embeddings
(nuss-and-bolts.com)
3 points
by
zanussbaum
1y ago
|
0 comments
3.
▲
The release of the first ModernBERT-based embedding models
(twitter.com)
1 points
by
zanussbaum
2y ago
|
1 comments
4.
▲
by
zanussbaum
2y ago
First embedding models trained from modern-bert-embed!
5.
▲
by
zanussbaum
2y ago
this was a huge inspiration for the post! i tried to highlight it in the blog but it might have gotten buried there are a few things that i wasn't able to figure out how to get access to/i wasn't sure if they were possible. f
6.
▲
by
zanussbaum
2y ago
oh in that case it was because i didn't know about them :) something to try next!
7.
▲
by
zanussbaum
2y ago
at least on my m2, the compiled kernel ends up using fast math anyways so using WGSL's fma didn't change anything about the actual kernel that gets run
8.
▲
by
zanussbaum
2y ago
thanks! and yes definitely not at CUDA levels :)
9.
▲
by
zanussbaum
2y ago
i tried using workgroup shared memory and found it slower than just recomputing everything in each thread although i may have been doing something dumb i'm excited to try subgroups though: https://developer.chrome.com/b
10.
▲
by
zanussbaum
2y ago
you're definitely right, 80% was a bit of an overestimation, especially with respect to CUDA it would be cool to see if there's some way to get better access to those lower-level primitives but would be surprised it does seem like
11.
▲
by
zanussbaum
2y ago
great question, to me webGPU sits a hair high level than CUDA or Vulkan. so you don't have the exact same level of control but can get to 80% performance of it without having to write different kernels specific to the hardware
12.
▲
Optimizing a WebGPU Matmul Kernel for 1 TFLOP
(zanussbaum.substack.com)
172 points
by
zanussbaum
2y ago
|
80 comments
13.
▲
by
zanussbaum
4y ago
Has been a huge boost over using Copilot. I accidentally was using Copilot instead of Codeium and was confused why the generations took so long until I realized! Great product
14.
▲
by
zanussbaum
6y ago
This made my day