6 ms·
There are Qwen3.5 27B quants in the range of 4 bits per weight, which fits into 16G of VRAM. The quality is comparable to Sonnet 4.0 from summer 2025. Inference
by throwdbaaway 7mo ago
There are Qwen3.5 27B quants in the range of 4 bits per weight, which fits into 16G of VRAM. The quality is comparable to Sonnet 4.0 from summer 2025. Inference speed is very good with ik_llama.cpp, and still decent with mainline llama.cpp.
- teaearlgraycold 7mo agoQwen3.5 35B A3B is much much faster and fits if you get a 3 bit version. How fast are you getting 27B to run? On my M3 Air w/ 24GB of memory 27B is 2 tok/s but 35B A3B is 14-22 tok/s which is actually usable.
- ece 7mo agoThe 27B is rated slightly higher for SWE-bench.
- throwdbaaway 7mo agoUsing ik_llama.cpp to run a 27B 4bpw quant on a RTX 3090, I get 1312 tok/s PP and 40.7 tok/s TG at zero context, dropping to 1009 tok/s PP and 36.2 tok/s TG at 40960 context. 35B A3B is faster but didn't do too well in my limited testing.
- ranger_danger 7mo agowith regular llama.cpp on a 3070ti I get 60tok/s TG with the 9B model, it's quite impressive.
- ranger_danger 7mo agoDon't sleep on the 9B version either, I get much faster speeds and can't tell any difference in quality. On my 3070ti I get ~60tok/s with it, and half that with the 35B-A3B.
- andai 7mo ago27B needs less memory and does better on benchmarks, but 35B-A3B seems to run roughly twice as fast.
- codemog 7mo agoCan someone explain how a 27B model (quantized no less) ever be comparable to a model like Sonnet 4.0 which is likely in the mid to high hundreds of billions of parameters? Is it really just more training data? I doubt it’s architecture improvements, or at the very least, I imagine any architecture improvements are marginal.
- otabdeveloper4 7mo agoThere's diminishing returns bigly when you increase parameter count. The sweet spot isn't in the "hundreds of billions" range, it's much lower than that. Anyways your perception of a model's "quality" is determined by careful post-training.
- zozbot234 7mo agoMore parameters improves general knowledge a lot, but you have to quantize more in order to fit in a given amount of memory, which if taken to extremes leads to erratic behavior. For casual chat use even Q2 models can be compelling, agentic use requires more regularization thus less quantized parameters and lowering the total amount to compensate.
- codemog 7mo agoInteresting. I see papers where researchers will finetune models in the 7 to 12b range and even beat or be competitive with frontier models. I wish I knew how this was possible, or had more intuition on such things. If anyone has paper recommendations, I’d appreciate it.
- zozbot234 7mo agoWith MoE models, if the complete weights for inactive experts almost fit in RAM you can set up mmap use and they will be streamed from disk when needed. There's obviously a slowdown but it is quite gradual, and even less relevant if you use fast storage.
- htrp 7mo agoany good packages you recommend for this?
- ljosifov 7mo agoSay more please if you can. How/why is ik_llama.cpp faster then mainline, for the 27B dense? I'd like to be able to run 27B dense faster on a 24GB vram gpu, and also on an M2 max.
- ac29 7mo agoik_llama.cpp was about 2x faster for CPU inference of Qwen3.5 versus mainline until yesterday. Mainline landed a PR that greatly increased speed for Qwen3.5, so now ik_llama.cpp is only 10% faster on token generation.