4 ms·
You can get those numbers with https://codeberg.org/ggz14/radiance-vllm-mxfp4 https://codeberg.org/ggz14/radiance-vllm-mxfp4. I also can get it on a single R970
by latentsea 24d ago
You can get those numbers with https://codeberg.org/ggz14/radiance-vllm-mxfp4 https://codeberg.org/ggz14/radiance-vllm-mxfp4. I also can get it on a single R9700 but the 75 ~ 80 t/s is only peak acceptance of very predictable tokens like coding or json, and averages lower for prose. It's still much faster than regular llama.cpp.
- julianlam 22d agoCan confirm. I've a single R9700 and have maxed out at 45 tokens/second on llama.cpp with Q4 Qwen 3.6 27B (with MTP)