Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
samhoss93
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
3 ms
·
1.
▲
by
samhoss93
4mo ago
piqc scans your Kubernetes cluster (Read-only) and identifies which models are running on the wrong GPU tier and what the cost attribution is. It runs in a minute. I'd like to hear the community's experiences/thoughts on ou
2.
▲
Show HN: Piqc – An open-source GPU waste scanner for LLM inference clusters
(github.com)
1 points
by
samhoss93
4mo ago
|
1 comments
3.
▲
Show HN: Piqc – GPU waste scanner for LLM inference clusters
(github.com)
3 points
by
samhoss93
4mo ago
|
0 comments
4.
▲
by
samhoss93
4mo ago
Agree. At high concurrency, you are better off spending the compute budget on parallel requests rather than draft prediction. The challenging part is that most deployment don't have static traffic profiles. A configuration that was ri
5.
▲
by
samhoss93
4mo ago
Great README. Genuinely one of the clearest walkthrough of inference internals. The KV cache section is worth lingering one as most of the OOM and throughput issues trace back to this and normally difficult to reason about. sequence leng