Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
philipturner
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
4 ms
·
1.
▲
by
philipturner
3y ago
To give more credit, I know your company was working on your own optimizations, which are prior art. It's possible that you made a 180 GB/s shader on your own (quite slow compared to my 319 GB/s). Or that the 319 GB/s wa
2.
▲
by
philipturner
3y ago
Yes. I look forward to a future where we don't have to wrestle between software and hardware. AI just runs, whether on CPU, GPU, TPU, QPU, or whatever architecture you want to design. Time to move on and focus on more important problem
3.
▲
by
philipturner
3y ago
The difference in bandwidth between M2 Max and data centers GPUs isn't that much (less than a factor of 5). The difference in compute is much, much larger. If you only have fast GEMV kernels, and not fast GEMM kernels, you're basi
4.
▲
by
philipturner
3y ago
For bandwidth-bound problems like large language models, you could also solve it with properly written CPU kernels (Mojo) and usage of the AMX accelerators for compute-intensive parts. I'd be more interested if they had a GPU port of S
5.
▲
by
philipturner
3y ago
antinucleon lol LLaMA is a memory-bound AI model, where the dominant factor in execution time is how fast the processor transfers weights from RAM to registers. LLaMA.cpp uses a misaligned memory pattern that's painful to the RAM I