3 ms·
1 4090, Qwen3.5-35B-A3B-UD-MXFP4_MOE, 64k context, 122 t/s. Llama.cpp
by instagib 7mo ago
1 4090, Qwen3.5-35B-A3B-UD-MXFP4_MOE, 64k context, 122 t/s.
Llama.cpp
- mirekrusin 7mo agoI believe it's mentioned that MXFP4 performs surprisingly bad, you may want to try other Q4s.