2 ms·
Btw, using llama.cpp you can achieve 55 - 45 token/second for processing/generation. if you use a qwen3.8-27B (IQ4_XS) verison, with decent quality in reasoning
by davada 14d ago
Btw, using llama.cpp you can achieve 55 - 45 token/second for processing/generation. if you use a qwen3.8-27B (IQ4_XS) verison, with decent quality in reasoning for coding/tool usage (with a 16GB nvidia).
I think now most of the struggle is getting a 32B-27B llm to work on 16GB/12GB card, because they are at least affordable/accessible for the time being compared to higher end models.
Recently the nvidia RTX 5090 32GB has reached price range of 7500 dollars (despite MSRP being around 2000 dollars when it was first launched). Crazy times.