Y
HN Search
Hacker News Search
new
|
comments
|
top
|
jobs
davada
searching PlanetScale…
1.
▲
2.
▲
3.
▲
4.
▲
5.
▲
6.
▲
26 ms
·
1.
▲
by
davada
7d ago
Btw, using llama.cpp you can achieve 55 - 45 token/second for processing/generation. if you use a qwen3.8-27B (IQ4_XS) verison, with decent quality in reasoning for coding/tool usage (with a 16GB nvidia). I think now most of
2.
▲
by
davada
7d ago
Its not so secretive, you can see the parameters they are using (arguments, and inference engine). I understand your point, but the only thing unusual/uncommon about it is the jargon used for LLM/AI systems.