3 ms·
Note that this is not the only way to run Qwen 3.5 397B on consumer devices, there are excellent ~2.5 BPW quants available that make it viable for 128G devices.
by tarruda 7mo ago
Note that this is not the only way to run Qwen 3.5 397B on consumer devices, there are excellent ~2.5 BPW quants available that make it viable for 128G devices.
I've had great success (~20 t/s) running it on a M1 Ultra with room for 256k context. Here are some lm-evaluation-harness results I ran against it:
mmlu: 87.86%
gpqa diamond: 82.32%
gsm8k: 86.43%
ifeval: 75.90%
More details of my experience:
- https://huggingface.co/ubergarm/Qwen3.5-397B-A17B-GGUF/discussions/8 https://huggingface.co/ubergarm/Qwen3.5-397B-A17B-GGUF/discu...
- https://huggingface.co/ubergarm/Qwen3.5-397B-A17B-GGUF/discussions/2 https://huggingface.co/ubergarm/Qwen3.5-397B-A17B-GGUF/discu...
- https://gist.github.com/simonw/67c754bbc0bc609a6caedee16fef89e8?permalink_comment_id=5991165#gistcomment-5991165 https://gist.github.com/simonw/67c754bbc0bc609a6caedee16fef8...
Overall an excellent model to have for offline inference.
- Aurornis 7mo agoThe method in this link is already using a 2-bit quant. They also reduced the number of experts per token from 10 to 4 which is another layer of quality degradation. In my experience the 2-bit quants can produce output to short prompts that makes sense but they aren’t useful for doing work with longer sessions. This project couldn’t even get useful JSON out of the model because it can’t produce the right token for quotes: > *2-bit quantization produces \name\ instead of "name" in JSON output, making tool calling unreliable.
- tarruda 7mo agoI can't say anything about the OP method, but I already tested the smol-IQ2_XS quant (which has 2.46 BPW) with the pi harness. I did not do a very long session because token generation and prompt processing gets very slow, but I think I worked for up to ~70k context and it maintained a lot of coherence in the session. IIRC the GPQA diamond is supposed to exercise long chains of thought and it scored exceptionally well with 82% (the original BF16 official number is 88%: https://huggingface.co/Qwen/Qwen3.5-397B-A17B https://huggingface.co/Qwen/Qwen3.5-397B-A17B). Note that not all quants are the same at a certain BPW. The smol-IQ2_XS quant I linked is pretty dynamic, with some tensors having q8_0 type, some q6_k and some q4_k (while the majority is iq2_xs). In my testing, this smol-IQ2_XS quant is the best available at this BPW range. Eventually I might try a more practical eval such as terminal bench.
- Aurornis 7mo ago> I did not do a very long session This is always the problem with the 2-bit and even 3-bit quants: They look promising in short sessions but then you try to do real work and realize they’re a waste of time. Running a smaller dense model like 27B produces better results than 2-bit quants of larger models in my experience.
- singpolyma3 7mo agoLots of people seem to use 4bit. Do you think that's worth it vs a smaller model in some cases?
- hnfong 7mo agoGenerally the perplexity charts indicate that quality drops significantly below 4-bit, so in that sense 4-bit is the sweet spot if you're resource constrained.
- Aurornis 7mo ago4 bit is as low as I like to go. There are KLD and perplexity tests that compare quantizations where you can see the curve of degradation, but perplexity and KLD numbers can be misleading compared to real world use where small errors compound over long sessions. In my anecdotal experience I’ve been happier with Q6 and dealing with the tradeoffs that come with it over Q4 for Qwen3.5 27B.
- amelius 7mo ago> This is always the problem with the 2-bit and even 3-bit quants: They look promising in short sessions but then you try to do real work and realize they’re a waste of time. It would be nice to see a scientific assessment of that statement.
- simonw 7mo agoThe project doesn't just use 2-bit - that was one of the formats they tried, but when that didn't give good tool calls they switched to 4-bit.
- tarruda 7mo agoIn my case it the 2.46BPW has been working flawless for tool calling, so I don't think 2-bit was the culprit for JSON failing. They did reduce the number of experts, so maybe that was it?
- stuaxo 7mo agoThere's at least one project they could use to repair the JSON and another that work takes a different approach.
- outlog 7mo agoWhat is power usage? maybe https://www.coconut-flavour.com/coconutbattery/ https://www.coconut-flavour.com/coconutbattery/ can tell you estimate?
- tarruda 7mo agoI don't think I've ever seen the M1 ultra GPU exceed 80w in asitop. Update: I just did a quick asitop test while inferencing and the GPU power was averaging at 53.55
- arjie 7mo agoWhat's the tok/s you get these days? Does it actually work well when you use more of that context? By the way, it's been a long time since I last saw your username. You're the guy who launched Neovim! Boy what a success. Definitely the Kickstarter/Bountysource I've been a tiny part of that had the best outcome. I use it every day.
- tarruda 7mo ago> What's the tok/s you get these days? I ran llama-bench a couple of weeks ago when there was a big speed improvement on llama.cpp (https://github.com/ggml-org/llama.cpp/pull/20361#issuecomment-4039467718 https://github.com/ggml-org/llama.cpp/pull/20361#issuecommen...): % llama-bench -m ~/ml-models/huggingface/ubergarm/Qwen3.5-397B-A17B-GGUF/smol-IQ2_XS/Qwen3.5-397B-A17B-smol-IQ2_XS-00001-of-00004.gguf -fa 1 -t 1 -ngl 99 -b 2048 -ub 2048 -d 0,10000,20000,30000,40000,50000,60000,70000,80000,90000,100000,150000,200000,250000 ggml_metal_device_init: tensor API disabled for pre-M5 and pre-A19 devices ggml_metal_library_init: using embedded metal library ggml_metal_library_init: loaded in 0.008 sec ggml_metal_rsets_init: creating a residency set collection (keep_alive = 180 s) ggml_metal_device_init: GPU name: MTL0 ggml_metal_device_init: GPU family: MTLGPUFamilyApple7 (1007) ggml_metal_device_init: GPU family: MTLGPUFamilyCommon3 (3003) ggml_metal_device_init: GPU family: MTLGPUFamilyMetal3 (5001) ggml_metal_device_init: simdgroup reduction = true ggml_metal_device_init: simdgroup matrix mul. = true ggml_metal_device_init: has unified memory = true ggml_metal_device_init: has bfloat = true ggml_metal_device_init: has tensor = false ggml_metal_device_init: use residency sets = true ggml_metal_device_init: use shared buffers = true ggml_metal_device_init: recommendedMaxWorkingSetSize = 134217.73 MB | ------------------------------ | ---------: | ---------: | ---------- | ------: | -------: | -: | --------------: | -------------------: | | qwen35moe 397B.A17B Q8_0 | 113.41 GiB | 396.35 B | MTL,BLAS | 1 | 2048 | 1 | pp512 | 189.67 ± 1.98 | | qwen35moe 397B.A17B Q8_0 | 113.41 GiB | 396.35 B | MTL,BLAS | 1 | 2048 | 1 | tg128 | 19.98 ± 0.01 | | qwen35moe 397B.A17B Q8_0 | 113.41 GiB | 396.35 B | MTL,BLAS | 1 | 2048 | 1 | pp512 @ d10000 | 168.92 ± 0.55 | | qwen35moe 397B.A17B Q8_0 | 113.41 GiB | 396.35 B | MTL,BLAS | 1 | 2048 | 1 | tg128 @ d10000 | 18.93 ± 0.02 | | qwen35moe 397B.A17B Q8_0 | 113.41 GiB | 396.35 B | MTL,BLAS | 1 | 2048 | 1 | pp512 @ d20000 | 152.42 ± 0.22 | | qwen35moe 397B.A17B Q8_0 | 113.41 GiB | 396.35 B | MTL,BLAS | 1 | 2048 | 1 | tg128 @ d20000 | 17.87 ± 0.01 | | qwen35moe 397B.A17B Q8_0 | 113.41 GiB | 396.35 B | MTL,BLAS | 1 | 2048 | 1 | pp512 @ d30000 | 139.37 ± 0.28 | | qwen35moe 397B.A17B Q8_0 | 113.41 GiB | 396.35 B | MTL,BLAS | 1 | 2048 | 1 | tg128 @ d30000 | 17.12 ± 0.01 | | qwen35moe 397B.A17B Q8_0 | 113.41 GiB | 396.35 B | MTL,BLAS | 1 | 2048 | 1 | pp512 @ d40000 | 128.38 ± 0.33 | | qwen35moe 397B.A17B Q8_0 | 113.41 GiB | 396.35 B | MTL,BLAS | 1 | 2048 | 1 | tg128 @ d40000 | 16.38 ± 0.00 | | qwen35moe 397B.A17B Q8_0 | 113.41 GiB | 396.35 B | MTL,BLAS | 1 | 2048 | 1 | pp512 @ d50000 | 118.07 ± 0.55 | | qwen35moe 397B.A17B Q8_0 | 113.41 GiB | 396.35 B | MTL,BLAS | 1 | 2048 | 1 | tg128 @ d50000 | 15.66 ± 0.00 | | qwen35moe 397B.A17B Q8_0 | 113.41 GiB | 396.35 B | MTL,BLAS | 1 | 2048 | 1 | pp512 @ d60000 | 108.44 ± 0.38 | | qwen35moe 397B.A17B Q8_0 | 113.41 GiB | 396.35 B | MTL,BLAS | 1 | 2048 | 1 | tg128 @ d60000 | 14.98 ± 0.01 | | qwen35moe 397B.A17B Q8_0 | 113.41 GiB | 396.35 B | MTL,BLAS | 1 | 2048 | 1 | pp512 @ d70000 | 98.85 ± 0.18 | | qwen35moe 397B.A17B Q8_0 | 113.41 GiB | 396.35 B | MTL,BLAS | 1 | 2048 | 1 | tg128 @ d70000 | 14.36 ± 0.00 | | qwen35moe 397B.A17B Q8_0 | 113.41 GiB | 396.35 B | MTL,BLAS | 1 | 2048 | 1 | pp512 @ d80000 | 91.39 ± 0.49 | | qwen35moe 397B.A17B Q8_0 | 113.41 GiB | 396.35 B | MTL,BLAS | 1 | 2048 | 1 | tg128 @ d80000 | 13.84 ± 0.00 | | qwen35moe 397B.A17B Q8_0 | 113.41 GiB | 396.35 B | MTL,BLAS | 1 | 2048 | 1 | pp512 @ d90000 | 85.76 ± 0.24 | | qwen35moe 397B.A17B Q8_0 | 113.41 GiB | 396.35 B | MTL,BLAS | 1 | 2048 | 1 | tg128 @ d90000 | 13.30 ± 0.00 | | qwen35moe 397B.A17B Q8_0 | 113.41 GiB | 396.35 B | MTL,BLAS | 1 | 2048 | 1 | pp512 @ d100000 | 80.19 ± 0.83 | | qwen35moe 397B.A17B Q8_0 | 113.41 GiB | 396.35 B | MTL,BLAS | 1 | 2048 | 1 | tg128 @ d100000 | 12.82 ± 0.00 | | qwen35moe 397B.A17B Q8_0 | 113.41 GiB | 396.35 B | MTL,BLAS | 1 | 2048 | 1 | pp512 @ d150000 | 54.46 ± 0.33 | | qwen35moe 397B.A17B Q8_0 | 113.41 GiB | 396.35 B | MTL,BLAS | 1 | 2048 | 1 | tg128 @ d150000 | 10.17 ± 0.09 | | qwen35moe 397B.A17B Q8_0 | 113.41 GiB | 396.35 B | MTL,BLAS | 1 | 2048 | 1 | pp512 @ d200000 | 47.05 ± 0.15 | | qwen35moe 397B.A17B Q8_0 | 113.41 GiB | 396.35 B | MTL,BLAS | 1 | 2048 | 1 | tg128 @ d200000 | 9.04 ± 0.02 | | qwen35moe 397B.A17B Q8_0 | 113.41 GiB | 396.35 B | MTL,BLAS | 1 | 2048 | 1 | pp512 @ d250000 | 40.71 ± 0.26 | | qwen35moe 397B.A17B Q8_0 | 113.41 GiB | 396.35 B | MTL,BLAS | 1 | 2048 | 1 | tg128 @ d250000 | 8.01 ± 0.02 | build: d28961d81 (8299) So it starts at 20 tps tg and 190 tps pp with empty context and ends at 8 tps tg and 40 tps pp with 250k prefill. I suspect that there are still a lot of optimizations to be implemented for Qwen 3.5 on llama.cpp, wouldn't be surprised to reach 25 tps in a few months. > You're the guy who launched Neovim! That's me ;D > I use it every day. So do I for the past 12 years! Though I admit in the past year I greatly reduced the amount of code I write by hand :/
- iwontberude 7mo agoThank you, I have been using way too much credits for my personal automation.
- woile 7mo agoJust a single m1 ultra?
- tarruda 7mo agoYes. Note that the only reason I acquired this device was to run LLMs, so I can dedicate its whole RAM to it. Probably not viable for a 128G device where you are actively using for other things.