2 ms·
with llama-cpp and offloading non-active experts (from MOE architecture) to cpu RAM, you can easily run 50 tok / s QWEN-3.6 35B on 8-12 GB of VRAM. KV cache is
by upboundspiral 4mo ago
with llama-cpp and offloading non-active experts (from MOE architecture) to cpu RAM, you can easily run 50 tok / s QWEN-3.6 35B on 8-12 GB of VRAM.
KV cache is a few GB, experts are ~3-5 GB (assuming q8 quant from Unsloth for example).
You can scroll through r/localllama and find tons of people getting useable speeds out of Qwen 35B.
24 tok / second on an ancient 1080ti
https://old.reddit.com/r/LocalLLaMA/comments/1tcc7h5/24_toks_from_30b_moe_models_on_an_old_gtx_1080_8/ https://old.reddit.com/r/LocalLLaMA/comments/1tcc7h5/24_toks...
100 tok / second on a 4070
https://old.reddit.com/r/LocalLLaMA/comments/1tjh7az/110_toks_with_12gb_vram_on_qwen36_35b_a3b_and_ik/ https://old.reddit.com/r/LocalLLaMA/comments/1tjh7az/110_tok...