31 ms·
Benchmarking Qwen3.8 27B quantizations: 4-bit holds up, 1-bit collapses
- sanjusangh 25d agoIsko ek karna hai
- purpleflame1257 25d agoThere's a real hole here at Q3. A critical breakpoint here is sub 16-GB cards, which covers the 5080, 5070 Ti, 5060ti, and several other cards from this generation and the last. It would be instructive to see where the quality knee is.
- jadbox 25d agoQ3 XL and Q3 XS are the two I'm trying to decide on
- dofm 25d agoYou might want to test this new dynamic GGUF: https://huggingface.co/ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF https://huggingface.co/ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF (I don’t know much about it, just saw a YouTube video about it last night)
- kennywinker 25d agoAnother one to try: https://huggingface.co/Jackrong/Qwopus3.8-27B-Flash-GGUF https://huggingface.co/Jackrong/Qwopus3.8-27B-Flash-GGUF Runs the 3bit model faster than the 2bit one runs on my old-ass card. Can’t vouch for its intelligence yet, but i suspect whatever loss in smarts it takes is made up for by the extra resolution.
- civvv 25d agoRunning Q3 on my AMD RX 9070XT. 32k context and 32/TPS. Apart from the context window preventing it from doing any large tasks, this thing is seriously powerful. I could probably push it to 64k context. Local open models are the future, and I am definitely getting a more powerful card. Very fun!
- slim 25d agoRunning Q3 on 5060ti with 64k context. It runs great
- Forgeties79 25d agoWhat are you offloading to ram (or even CPU)? I’m using a 9080 (not XT) and having trouble with context/token rates
- civvv 25d agoI’m running Qwen3.8-27B-Unleashed UD-Q3_K_XL, which is a ~12.3 GiB Q3 quant, fully offloaded to the 16 GB 9070 XT. I disabled the vision projector to save VRAM and use one inference slot, Flash Attention, Q4 KV cache, --fit off, and --ctx-checkpoints 0. I’m running it with a 64K context window. The AMD driver also needs to be recent enough for ROCm 7.14; I targeted Adrenalin 26.6.4 or newer.
- 7speter 25d agoYou can offload the vision projector to CPU/sysRAM
- brynx97 25d agoCould you comment more on how you set this up? I have a mostly idle 9070XT I use for gaming, and I was considering using it with the newer local open models. Many thanks.
- dofm 25d agoThere is an interesting new dynamic 3 bit quantisation I have been meaning to test: https://huggingface.co/ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF https://huggingface.co/ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF Luke of Luke’s Dev Lab on YouTube had a look at it. It seems to outperform the typical 3-bit quantisation but whether it outperforms the new Unsloth dynamic I don’t know.
- selectodude 25d agoI have a 5080, three OpenAI Pro token resets, and I’m on paternity leave. Astra seems pretty clever. Maybe I’ll give it a task.
- deleted 25d ago[deleted]
- spider-mario 25d ago> Second, besides noise (bars are Wilson 95% confidence intervals, very conservative for run-to-run noise), there is little difference down to 4-bit; only the 2-bit scores a bit lower. Confidence intervals have nothing to do with run-to-run variation. They have little to do with anything people usually ascribe to them (https://link.springer.com/article/10.3758/s13423-015-0947-8 https://link.springer.com/article/10.3758/s13423-015-0947-8 ), but even less with run-to-run variation (https://link.springer.com/article/10.1007/s10654-016-0149-3 https://link.springer.com/article/10.1007/s10654-016-0149-3 misconception 22).
- jnwatson 25d agoMind blown. The more I read about statistics, the less I know.
- exogenousdata 25d ago“There are three kinds of lies: Lies, damned lies and statistics.” - Mark Twain (attributed but unsubstantiated to Benjamin Disraeli)
- fr2029 25d agothe 2nd derivate of shannon covariance of noise begs to differ
- deleted 25d ago[deleted]
- stared 25d agoPoint taken, but there is a much more fundamental issue with it - and precisely why I wrote "very conservative". It is a different problem if we pick two sets from the same data distribution, A and B, and first we have a score on A, then on B. Here we re-run on precisely the same set of Terminal Bench 2.1 problems. It may be that results are so random between runs that each single task has the same probability in a Bernoulli distribution. But more likely, many problems are easy (i.e. each run will solve them consistently), many are too hard (i.e. no run is going to solve them) and only a fraction is somehow in between. Maybe there is some good trick to find a proper distribution, but to my knowledge, we would need to run it at least two times on TB2.1 to get any more educated estimates. That said, I am open to new ideas. That said, I consider frequentist probability a dirty trick, and that Bayesian is the proper way of doing things (vide David J.C. MacKay" Information Theory, Inference, and Learning Algorithms" and Cam Davidson-Pilon "Probabilistic Programming & Bayesian Methods for Hackers" https://www.inference.org.uk/itprnn/book.pdf https://www.inference.org.uk/itprnn/book.pdf, https://dataorigami.net/Probabilistic-Programming-and-Bayesian-Methods-for-Hackers/ https://dataorigami.net/Probabilistic-Programming-and-Bayesi...).
- bellowsgulch 25d agoQwen3.8 27B seems like it was clearly supposed to be a high-end consumer open-weights model, but the t/s is so low for me on my old M1 Max 64GB that I hope others are getting use out of it. Unfortunately, the calculus has changed and it seems cheaper to me to just use MiMo V2.5 for pennies or DeepSeek V4 Flash instead of using Qwen anymore unless I need a local model specifically for doing reverse engineering work that gets otherwise rejected.
- spider-mario 25d ago> Qwen3.8 27B seems like it was clearly supposed to be a high-end consumer open-weights model, but the t/s is so low for me on my old M1 Max 64GB that I hope others are getting use out of it. Have you tried it with MTPLX? I get around 30 tok/s with it, also on an M1 Max with 64GB.
- Xeoncross 25d agoNice, which model quantization is this? Is it on huggingface?
- spider-mario 25d agoMTPLX is this software: https://github.com/youssofal/MTPLX https://github.com/youssofal/MTPLX I tried it with the author’s 4-bit quant of Qwen 3.8 27B: https://huggingface.co/Youssofal/Qwen3.8-27B-MTPLX-Optimized-Speed https://huggingface.co/Youssofal/Qwen3.8-27B-MTPLX-Optimized... (but no need to download it manually; MTPLX will ask which one you want).
- SwellJoe 25d agoEven at 30 t/s, 3.8 thinks so long, even on medium, it still takes 3x or more longer than any cloud model, in my testing.
- lowbloodsugar 25d agoI've got an M1 Max 64GB too. It's just not an LLM-class workstation. Give it a year and buy an M7 and you'll be laughing. Right now is a really bad time to invest in anything - using the cloud is the cheapest option, especially for open weight models.
- deleted 25d ago[deleted]
- InvectusXIV 25d ago[flagged]
- quietraster 25d agothe 4-bit matching bf16 on terminal-bench is a useful data
- dvh 25d agoCould this be used to estimate how many fingers LLM have?
- zrail 25d agoI've been running Unsloth IQ3_S on my 5060ti with mmproj offloaded, getting 600-1000 prefill and 30-50 tg with this config: /data/llm/llama.cpp/build/bin/llama-server --threads 4 --threads-batch 8 --batch-size 4096 --ubatch-size 256 --port 9999 --temp "1.0" --top-p "0.95" --top-k "20" --min-p "0.0" --presence-penalty "0.0" --reasoning auto --reasoning-preserve --reasoning-budget 4096 --gpu-layers-draft all --spec-type draft-mtp,ngram-map-k4v,ngram-mod --spec-draft-n-max 3 --spec-draft-p-min 0.75 --spec-ngram-mod-n-match 24 --spec-ngram-mod-n-min 4 --spec-ngram-mod-n-max 16 --spec-ngram-map-k4v-size-n 8 --spec-ngram-map-k4v-size-m 16 --spec-ngram-map-k4v-min-hits 1 --n-gpu-layers all --ctx-size 131072 --repeat-penalty 1.0 --jinja --metrics --model /data/llm/models/unsloth/Qwen3.8-27B-UD-IQ3_S.gguf --chat-template-file /data/llm/models/qwen3.6-chat-template.jinja --fit off --flash-attn on --cors-origins localhost --mmproj /data/llm/models/unsloth/Qwen3.8/mmproj-BF16.gguf --no-mmproj-offload --parallel 1 --kv-unified --cache-type-k q4_0 --cache-type-v q4_0 --cache-type-k-draft q4_0 --cache-type-v-draft q4_0
- zrail 25d agoToo late to edit, but a few other things to note: I minmaxed the draft config. On my typical coding workloads it gets around 70% acceptance, more variable on prose. The chat template is froggeric's fixed qwen template, v22.5 as of today.
- Farmadupe 25d agohmm, assuming that this article is part written by claude and part human-written, can anyone help me find a rule of thumb for "how to know if the article is worth reading"? Because on the one hand, the prose and the presentation is painful (narrating irrelevant points, nonlinear X-axes, ambiguous chart labels, etc etc), But on the other hand, the result that I'm assuming the author means to communicate ("on these evals, generation quality seems fairly good") sounds worthwhile to share? Because I really struggle with this question at the moment. Am I allowed to draw an adverse inference that "if the writeup presents irrelevant text side by side with the data, then this may be a sign that the author does not understand the task that they are attempting to write up"?
- clircle 25d agoI think the advice is the same regardless of AI use: read articles written by authors that have a history of high quality writing.
- cogman10 25d agoIMO, whether or not an LLM was used in the writing process doesn't really matter and I think it's a bit annoying that articles are being dismissed out of hand because of that. The line is "Is this an interesting and accurate article that concisely makes it's case". LLMs love to burn paragraphs writing about nothing which is why it's generally poor writing. Humans can do the same thing if they are trying to make very little information feel more substantial. I say, stop trying to determine if an LLM was used and start judging based on your subjective measure that you'd have used before LLMs became widespread.
- dotinvictim 25d agolocal llm don't make sense currently consumer compute is not upto mark it may take atleast 7 more years to be usable
- kennywinker 25d agoIt literally is usable now. A 5060 for $800 can run qwen3.8-27b 4bit at >40t/s, and the model beats opus 4.6 (max).
- TomBombadildoze 25d agoBeats Opus 4.6 at what exactly? It certainly isn't code. I use a combination of a Claude Max subscription and local inference, including qwen3.8-27b, 4bit. I have found qwen to be absolutely useless at anything but very specific, surgical code changes. In my experience, for anything even remotely nuanced, a frontier model is required.
- brandon272 25d agoWhat configuration? What harness? These matter greatly to how a local model performs, in my experience.
- kennywinker 25d agohttps://artificialanalysis.ai/?models=qwen3-8-27b%2Cclaude-opus-4-6-adaptive#artificial-analysis-intelligence-index https://artificialanalysis.ai/?models=qwen3-8-27b%2Cclaude-o... Index methodologies here: https://artificialanalysis.ai/evaluations/artificial-analysis-intelligence-index https://artificialanalysis.ai/evaluations/artificial-analysi... Also see some specific benchmarks here: https://huggingface.co/Qwen/Qwen3.8-27B https://huggingface.co/Qwen/Qwen3.8-27B e.g. qwen scores 61.7 on swe bench pro, while opus 4.6 scores 53.4. If you want to argue with the benchmarks, go for it. Fwiw i am not saying qwen3.8-27b is better or as good as the frontier. But i am saying it has crossed the threshold and is now a useful tool for coding and debugging. From my experience, Qwen3.6-35b-a3b was what you describe - it could do surgical edits only.
- 25d ago
- john_rood 25d ago[flagged]
- kouteiheika 25d agoNote that these quants are not quantized uniformly, so 4-bit isn't actually a "true" 4-bit here, so these observations won't necessarily hold up to other quants which might be done differently.
- wgd 25d agoIt looks like they tested Q4_K_M which should be just the standard K-quant without any imatrix calibration. The smaller ones are indeed dynamic though.
- deleted 25d ago[deleted]
- syntaxing 25d agoI’m more curious how each 4 bit quant compares. It seems like NVFP4 outperforms Q4_K_M in terms of speed and top 1 but is only good for expensive Nvidia cards
- sharmajai 25d agoThis confirms a theory I have to explain the minimal loss in quality when using lower quants (I use IQ3_XXS with an 8-bit KV cache) and the XHIGH (default) thinking level. It's well-known that while quantization affects the sampling probability distribution (given the same context, which next token is the most probable), Qwen 3.8 27b seems to offset that by just thinking more and as a result eventually finishing the task (benchmark or otherwise). So as long as the thinking (albeit longer) is sound, this leads to the same success rate (as shown in the article) but potentially at the cost of more tokens and hence more time. I think it'll be further useful to chart each quantization's used tokens as well, in addition to the success rate. Thanks for doing and sharing the research!
- seemaze 25d agoAs they say, time is money. In the age of the rampocalypse, the peasants may not have a choice between the two.. time it is!
- conmod278 25d agoComputer science has known the tradeoffs between memory and compute since ages ago. The same could be reflected here.
- chmod775 25d agoSmaller models are also generally faster, so thinking "more" may not matter and may even come out ahead.
- celrod 25d agoIf Q4 takes less than 1.3x as many tokens as bf16 or q8, it could still end up being faster, given how decode tends to be bandwidth bound. The kv cache was still bf16, so a few ops are the same between quants.
- anon291 25d agoI personally think thinking is basically variable but rate precision. If you are in a 4bit mode but need 2x as many tokens you're just doing fp8 with hoops( of course 4bit multiply is faster)
- rvba 25d agoThose benchmarks are very interesting. But is there any model that actually works in a decent way at quantization of 1?
- alentred 25d agoI would be very interested in a similar benchmark for *KV cache* quantizations. I use Qwen3.8 27B Q4_K_M for coding sometimes and therefore need a relatively long context. I settled on q8_0 because it is the only way to fit the model + 100k tokens into 24GB VRAM, but still wonder what am I loosing in quality, and what other options are there. I also heard that KV cache quantization matters more with longer contexts. It may be interesting to benchmark this too: what the quality looks like on different combinations of model quantization × KV cache quantization × context size.
- quotemstr 25d agoYou don't have to quantize all layers and all dimensions uniformly, FWIW
- skolos 25d agoOn many models that I tested in past context quantization had very bad effect on model performance. However qwen3.8 27b is different. I'm now running NVFP4 quantized both weight and cache on my RTX5090 and getting excellent results: 264k cache allocated for pool, 10k tok/s prompt processing, 200 tok/s generation for single stream, or 801 tok/s generation for 8 concurrent streams. Also have about 2Gb vram left for use of OS. my coding agents regularly reach 200k context used without noticeable degradation. P.S. I used setup from: https://github.com/seanyourhighness/vllm-sm12x-nvfp4-dflash2 https://github.com/seanyourhighness/vllm-sm12x-nvfp4-dflash2
- D13Fd 25d agoI’m running 27B on a 5090 as well, and the results have been really strong. It does almost as well as, and sometimes better than, a 121gb DS4 model running on an M5 Max 128gb. 27B also flies on the 5090, and at medium think it returns results many times faster than my DS4 setup (the default xhigh is basically broken, though). For the kinds of things I use a local model for (legal document review), it’s just spectacular. It also has good vision support. I’ve been using 27B more and more over DS4.
- redox99 25d agoThose are really nice numbers. With that t/s, no network latency or queueing it must feel much snappier than cloud models.
- anyfoo 25d agoI have a very interesting self-made coding benchmark, very intricate and technical, but 100% a real world problem I had to solve. I’m not going to further elaborate, since I don’t want future models to train on the solution. To my own surprise, Q6_K_XL (from unsloth) comes up with a solution, anything Q5 doesn’t. To further surprise me, so far only the XL Q6 variant managed to solve it. The problem, at least as stated, seems to be right on the edge of what the Q6 quantization can do. Unfortunately even a successful run is rather long, so I don’t have a whole lot of data. But the whole thing sure made me doubt the common idea that you wouldn’t perceive a difference until crossing past 4 bits quantization.
- teaearlgraycold 25d agoI thought the wisdom is more so don’t bother going below 4bit and you won’t see a difference above 8bit.
- anyfoo 25d agoDepends on the actual audience, I guess. My stated “wisdom” comes in part from /r/LocalLLaMa, and my impression is that the tasks that users there give their models to try them out lean towards rather simplistic, on the reasoning side. But there I literally did read “you don’t need anything better than 4 bpw” a bunch of times.
- neonstatic 25d ago> my impression is that the tasks that users there give their models to try them out lean towards rather simplistic, on the reasoning side. That sub is games, porn, and complaints about not having money. A waste of time.
- anyfoo 25d ago> As you may see, the scores are around the random guessing level, with the smallest model being below that threshold. Err… can someone explain to me what is meant here? Surely the model wouldn’t consistently “guess wrong” compared to randomly, as that would be better. I guess some things like general coherency (i.e. is it even readable or gibberish) factor into that score?
- magnat 25d agoThose are multiple-choice questions. If some of them are "trick questions", where obvious answer (e.g. the value taken directly from question's text) is wrong, bad model might perform worse than a dice. On the other hand, not sure where from 25% baseline for random answers come from. Since this is multiple-choice-out-of-4 test, random guessing should be correct in 1 in 15 cases, not 1 in 4.
- mrbonner 25d agoI use a 2-bit quant from Unsloth on my MBP M5 32GB of RAM. It run slower than molasses at 2 too/s kind of thing. Not sure it is usable at that rate for anything.
- seamossfet 25d agoIf you want to do a 1-bit model you have to QAT at pre-training with way more data than chinchilla to compensate for the cliffs (like 50x). Quantization on an existing pre-trained model will almost always collapse at 1-bit
- KennyBlanken 25d agoIt's strange that the author has completely ignored the 3 bit quants which allow someone with a 16GB GPU to have 100-120k and still get full performance. You can't run any of the 4-bit quants on a 16GB gpu with enough context to be useful for all but the most basic tasks. General purpose agents can need up to 30k just to reply with "1+1=2" because their prompting is so overloaded. 60-70k is decently usable, still not great for anything complex. A long running task in a general purpose agent can easily hit 100k. What the vast majority of people care about is performance around what desktop consumer GPUs can run. 8GB, 10, 12, and 16GB of VRAM. What do models that will run at full performance, do? Also important to know is how Qwen3.8-27B stacks up against qwen3.6-35B-A3B, which due to being MoE, will run on a 16GB card with plenty of speed 90% of the time, at higher quant - so you get more parameters and better quant. But 3.8 is supposed to be "better", so...?
- nozzlegear 25d ago1-bit, 2-bit, penny and dime.
- kmike84 25d agoMeasuring quality e2e definitely makes sense. But I think there is a bit more to this: > Measuring token prediction differences (KL-divergence, top-1 predictions) is easy, but it does not tell us whether the model gets worse at solving tasks. A common issue is that it's rarely mentioned on which dataset KL-divergence is computed. It seems the most common dataset is wikitext - maybe because of tradition, to make numbers more comparable? It's not measuring how well the model follows bf16 on agentic tasks. I've been trying to check KLD recently for some quants of Qwen 3.8 27B, and the numbers are dramatically different, depending on which dataset you use. KLD computed on agentic traces is much higher, and top-1 % is way lower than if you compute it on chats or wiki text. You look at a published number, and see "oh, nice, top1 is 99% - quant is different just in 1 token out of 100", but chances are it's computed on wiki, and on agentic / coding it can be 10 tokens out of 100. Common intuition is that on agentic tasks errors compound, and that's why it degrades more than metrics show - but maybe the metrics themselves are also wrong, too optimistic. Still investigating it though :)
- xscott 25d ago> A common issue is that it's rarely mentioned on which dataset KL-divergence is computed. It seems the most common dataset is wikitext Thank you for calling this out. Using Wikipedia snippets for these is a terrible choice. I did a bunch of KL and other stats with the five Gemma 4 models, and the results were non-obvious. Anthropomorphizing: Gemma 4 31B: "I guess we'll pretend I said this, but it's not me." (Baseline for stats) Gemma 4 26B: "Dude, I'm certain I wouldn't have said this." (Bad KL) Gemma 4 12B: "Umm, Me either!" (Similarly Bad KL) Gemma 4 E4B: "I might say almost anything, this is fine." (Much better KL!!!) Gemma 4 E2B: "I'm basically a toy. Let's play a game!" (Same KL as E4B) Anyway, for comparing quantizations, it seems like the largest precision version should be given a one-shot prompt, and the result from that should be used as the corpus for the quantized versions.
- nullc 25d ago> Anyway, for comparing quantizations, it seems like the largest precision version should be given a one-shot prompt, and the result from that should be used as the corpus for the quantized versions. Way better than wikitext-- but tells you nothing about errors tending to compound or cancel out. Like say a test shows that only one token in a 10,000 token test would be different. Sounds very close, ship it!-- but what if trajectories with that single different token guarantees failure because it sets in motion a cascade of differences that ultimately result in a final distribution that doesn't include the solution?
- qsbuilder 25d agoI always wonder is it safe to run one of these models on a personal pc, or do you guys recommend something like docker, sorry a bit new to all of this.
- bitwize 25d agoYes, it's fine to run a model on bare metal. The model is just a token predictor. Leave out the fine semantics about this; it's a function taking a set of input tokens to output tokens. It can't mess with your computer or files until you hook it to a harness, which interprets some of the model's output as commands to execute. So, model on bare metal, harness in a container or VM.
- paidx 25d ago[flagged]
- hefu_hk 25d ago[flagged]
- Neat_comfort007 25d ago[flagged]
- nateb2022 24d ago> For comparison, DeepSeek V4 Flash 0731, a 284B model, costs around $0.1/Mtok for output from the cheapest providers on OpenRouter. I am not sure how much of this difference comes from the efficiency of running models at scale, pricing strategy, or popularity. Due in large part to DeepSeek's MoE architecture. Generating 1 token through DS4's 13B active params requires roughly half the FLOPs to generate a token through the dense Qwen 27B.