Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
brausepulver
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
by
brausepulver
8d ago
As far as I understand: 1) it's very fast (they claim 40-200x faster than frontier models [1], would roughly line up with it doing diffusion) 2) each answer carries a calibrated probability (ie. frequency of outcome is close to predict
2.
▲
by
brausepulver
17d ago
You need to separate memory capacity and bandwidth. Looping decreases memory capacity/FLOP but not bytes loaded/FLOP, since weights need to be loaded again for the 2nd pass. Plus (depending on the method used) capacity required fo
3.
▲
by
brausepulver
1mo ago
What LLM-specific hardware improvements should one expect? Seems to me that LLM inference is simple architecturally (matmul et al) so most scaling in hardware should come from general improvements (memory BW, packaging, interconnect, power)
4.
▲
by
brausepulver
2mo ago
Consider that a lot of the resistance to just existing is external expectations. I feel like I should always be doing something because I'm pressured to, or that small things bother me because I feel I need to perform. Being content