3 ms·
No way you can get that much each without extremely quants For a single r9700 you have 637 GB/s and for qwen 3.8 27b q4_k_xl the maximum tg/s is 33 before mtp
by pyrolistical 27d ago
No way you can get that much each without extremely quants
For a single r9700 you have 637 GB/s and for qwen 3.8 27b q4_k_xl the maximum tg/s is 33 before mtp
Now if you meant 4xr9700 tensor parallelism with mtp, 80 tg/s starts to make sense
- naasking 27d agoQ4 isn't an extreme quant, and I average 75 toks/s on code, 45 tok/s on prose with MTP.
- latentsea 27d agoYou can get those numbers with https://codeberg.org/ggz14/radiance-vllm-mxfp4 https://codeberg.org/ggz14/radiance-vllm-mxfp4. I also can get it on a single R9700 but the 75 ~ 80 t/s is only peak acceptance of very predictable tokens like coding or json, and averages lower for prose. It's still much faster than regular llama.cpp.
- julianlam 24d agoCan confirm. I've a single R9700 and have maxed out at 45 tokens/second on llama.cpp with Q4 Qwen 3.6 27B (with MTP)