3 ms·
I'm running unsloth/Qwen3.6-35B-A3B-UD-Q8_K_XL on an M3 Max, 64GB at ~57 t/s with llama-server
by juancn 5mo ago
I'm running unsloth/Qwen3.6-35B-A3B-UD-Q8_K_XL on an M3 Max, 64GB at ~57 t/s with llama-server
- brcmthrowaway 5mo agoPrefill speed and 27B number?
- juancn 5mo agoPrefill is around ~600 t/s. I don't remember what the 27B was, I tried a 27B with different quantization at some point for that one, but I settled on the 31B.