Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
pidtom
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
by
pidtom
6mo ago
I built TurboQuant+ ( https://github.com/TheTom/llama-cpp-turboquant ), the llama.cpp implementation of this paper with extensions: asymmetric K/V compression, boundary layer protection, sparse V dequant, and this w
2.
▲
by
pidtom
6mo ago
Hey that's me! AMA
3.
▲
Skipping 90% of KV dequant work speeds up LLM decode by 22%
(github.com)
1 points
by
pidtom
6mo ago
|
0 comments