22 ms·
Qwen 3.8 27B
- brcmthrowaway 2mo agoMy Strix Halo is about to go overdrive!
- tosh 2mo agoalso cool: Qwen 3.8 27b is multi modal!
- gurkwart 2mo agostrong visual reasoning apparently, which is nice. still lacking native audio however. hoping for more companies to embrace the spirit of something like `gemma-4-12b-qat` for actual multi-modality (text, image, video, audio).
- TomGarden 2mo agoReally excited to see what people do with this. 3.7 27B was probably the best compromise between size and intelligence to run on consumer hardware
- geek_at 2mo agoDo you mean 3.6 27b? Because qwen 3.7 didn't have an open weight version
- erdaltoprak 2mo agoThis is one of the most important model releases since most use cases don't need SOTA/Frontier If you want Qwen3.8-27B Serving Configs for the DGX Spark vLLM NVFP4 and RTX 4090 llama.cpp GGUF I added the setups here https://x.com/ErdalToprak/status/2088299678085308761?s=20 https://x.com/ErdalToprak/status/2088299678085308761?s=20
- water-drummer 2mo agoFor those who don't wanna open twitter: https://xcancel.com/ErdalToprak/status/2088299678085308761?s=20 https://xcancel.com/ErdalToprak/status/2088299678085308761?s...
- chvid 2mo agoThese are massive improvements - and something you can actually run on a laptop.
- deleted 2mo ago[deleted]
- tosh 2mo ago27b dense model at Opus 4.6 level Opus at home I hope there also will be a new ~10b variant
- yassa9 2mo agocan you tell me ideas of usecases of 9 or 10B language models ? I cant find any usecases other than training a lora on them to give good bash commands for example
- tosh 2mo agothey are all overlapping but: categorization, information retrieval, semantic search, image description also with the model as part of an agentic system with tool calling (edit: it is quite impressive what a small model in a feedback loop can do)
- mring33621 2mo ago9B Qwen models are good and fast for local python coding tasks.
- SwellJoe 2mo agoI use Gemma 4 12B in the 4-bit quantization for all sorts of vision tasks (image sorting, classification, description). It's also good for the same sorts of things for text (but there are probably better/faster models for text, 12B just happens to excel at vision tasks). The Qwen 9B is also very good for those tasks. If you need to do any kind of "search the web, grab some data, do some kind of action" tasks, these small models are perfect for that. Scraping data in a fuzzy format into a database or report or spreadsheet, producing a dashboard of news, etc. Small local models can also be used for sub-agent tasks in most agent harnesses. But I'd probably run a larger MoE for that; they're faster and have broader knowledge. The dense models, even very small ones, are not blazing fast. I don't code with any models small enough to run locally, at least not so far. Qwen 3.8 27B might be the tipping point, though. It's looking really promising, though it's probably slow enough that I won't ever actually use it. I'd rather pay $100/month for a faster model, even if Qwen 3.8 turns out to be smart enough for most of my work. Running it locally with the 8-bit quantization is going at 12-30 t/s, depending on how much context it's chewing on. So, if all you do with AI is coding, then you're better off doing it in the cloud. But, there's lots of things a small model can do that aren't coding.
- scrlk 2mo agoBeats Opus 4.7 Max (w/ Claude Code) on DeepSWE (42.2 vs 40). Looks like Qwen's 27B models continue to pack some punch. Unsloth's GGUF quants are up: https://huggingface.co/unsloth/Qwen3.8-27B-GGUF https://huggingface.co/unsloth/Qwen3.8-27B-GGUF
- Foobar8568 2mo agoConsidering the clusterfuck that is opus 5 or even fable, if Qwen 27B is trully better than Opus 4.7 Max, I will rejoice.
- UncleOxidant 2mo agoIf it's as good as Sonnet 4.6 for most things I'd be happy.
- ranguna 2mo agoSame
- ferrouswheel 2mo agoYeah Opus 5 is almost unusable as a daily driver without making me go insane from excessive claude babble.
- edg5000 2mo agoThat's crazy, considering the massive size difference. But the small Qwen models are known for punching above their weight.
- nblgbg 2mo agoIs there any advantage to using the model from Unsloth compared with https://huggingface.co/Qwen/Qwen3.8-27B-FP8 https://huggingface.co/Qwen/Qwen3.8-27B-FP8 ?
- benxh 2mo agoDepends on what software/hardware you'll run it. GGUFs from Unsloth can run on pretty much every single potato; full weights need beefy gpus
- kristopolous 2mo agoq4km is about 48 tps on a 4090. my llama.cpp params are --flash-attn on --parallel 1 --load-mode mmap
- m_ke 2mo agoWith spec decode should easily get to >100tps on my dual 3090s qwen 3.5 27b was running at around 110tps using the config from https://github.com/noonghunna/club-3090 https://github.com/noonghunna/club-3090 make that 200tps on a single 5090, 4x faster than opus https://x.com/radixark/status/2088285681131110446 https://x.com/radixark/status/2088285681131110446 devs about to get handed a two 5090 box each and told to max that out
- lta 2mo agoI've been experimenting with a few settings in my 4090 , and if 3.6 run at 90-110 tps, 3.8 staya below 80 tps and it's most often at 60 tps. I'm using flash attention, mtp speculative decoding (n=2). I've looked at the club-3090 repo, but haven't found anything meaningful but get back to 3.6 performance
- KronisLV 2mo agoI hope really badly that we'll get a new 35B A3B or similar MoE model! I also miss the Qwen 3 Coder Next, which was 80B A3B, there are quite a few use cases where a non-dense model <100B would be the sweet spot (when you have the VRAM but not the TDP or compute power). Heck, I'd gladly take A5B or A8B or even A10B as a sort of middle ground. Also alternate link for viewing the images without signing in: https://xcancel.com/Alibaba_Qwen/status/2088280182356611304 https://xcancel.com/Alibaba_Qwen/status/2088280182356611304
- Casteil 2mo agoI'm hoping too that they'll put out some MoE variants. Qwen3.5:122b:a10b can run about twice as fast as this 27b dense model. Edit: Like its predecessors, 3.8 seems really inclined to overthinking, and on a 27b dense model that's kind of painful. I think I'm going to stick with gemma4:26b-a3b as my go-to because it runs about 4x as fast and tends to only need a fraction of the tokens in its 'thinking' stage to get the same or similar answer.
- Phemist 2mo agoDid you try the claude reasoning traces finetune for qwen3.6? I find that it works muuch better. I assume the same 3.8 finetune will be released at some pointas well. Edit: link - https://huggingface.co/rico03/Qwen3.6-27B-Claude-Opus-Reasoning-Distilled https://huggingface.co/rico03/Qwen3.6-27B-Claude-Opus-Reason...
- satvikpendem 2mo agoReduce or turn down thinking: https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
- Casteil 2mo agoYeah, that's probably the answer given that it apparently defaults to 'xhigh'.
- altruios 2mo agoremember to let llama.cpp catch up to anything new in this model. Save your judgment until about 2 weeks of use.
- chrismartin 2mo ago'Good' news, there seems to be nothing new architecture-wise. Same as Qwen 3.5 and 3.6, so llama.cpp doesn't know the difference.
- expedited123 2mo agoKinda was expecting to see Gemma 4 26B in benchmark comparisons :(
- kamranjon 2mo agoSince Qwen 3.6 27b outperforms Gemma 4 26b in most benchmarks I'm not sure the value - also Gemma 26b is a MOE model whereas this is a dense model, so not typically direct competitors at their sizes - Gemma 4 31b comparison would be interesting though.
- expedited123 2mo agoI see! Thanks.
- anana_ 2mo agoMonstrous benchmarks! Hoping it is not benchmaxxed.
- sheepscreek 2mo agoI thought the same. But why claim something so shocking when it can easily be discredited and puts your reputation at risk? If they’re claiming Opus 4.6 level, I expect it to at least match Sonnet 4.6.
- kunver 2mo agoWelcome deepseek flash flash!
- deleted 2mo ago[deleted]
- NorwegianDude 2mo agoIf the benchmarks are a real indication, we now have a local model that is runnable on a high-end personal PC that trades blows with the leading model Claude Opus 4.6 Max from half a year ago. Insane if that is the case. Downloading now!
- throwaway613746 2mo ago[dead]
- onlyrealcuzzo 2mo agoIf the benchmarks don't lie, this is getting very close to Opus 4.6 capability - which was the turning point for me for when AI was "good enough" that it became very hard to justify not using it. I'm sure there's some benchmaxxing going on, and some things you get only with a a larger model. But I'm feeling pretty confident if not by Gemma 5 than by mid 2028 we'll have local models that are almost always as good as Opus 4.6 was and in many cases far better.
- DanielHB 2mo agoWhat kind of things you only get with a larger model?
- versteegen 2mo agoIME using 5.6 Luna and DS V4 Flash, I notice that although they are excellent at programming, even Opus-like in the way they try to debug, the thing they are worst at is inferring user intent and making good decisions with little information. They are absolutely terrible at that, will misinterpret small wording ambiguities. I suspect that's an ability you can't add with RL training, that it requires the depth of understanding from vast pre-training.
- johnnyApplePRNG 2mo agoSo add a pre-ingestion agent to your Pi Coding Agent subagent repertoire, problem solved.
- NamlchakKhandro 2mo agoWoah woah woah buddy... suggesting that people use anything other than ClaudeCode or Codex is simply not allowed around these parts.
- solarengineer 2mo agoYour negative impression is surprising. Could you point to any threads to support your message?
- kunver 2mo agoLooks like a pretty significant improvement on the DeepSWE benchmark compared to the previous 27B model.
- alpha_trion 2mo agoNICE, i've been waiting for this drop, thanks for posting this
- TomGarden 2mo agoAny tips on the best approach at running this at an M4 Max 128GB? Token throughput was a bit slow with the last 27B one (MLX), ended up using the A3B variant but if I could get this one to reasonable speed I'd much prefer it.
- LoganDark 2mo agoUnfortunately, that chip just doesn't really have the memory bandwidth to run this (or nearly any) model at acceptable speeds. I have the exact same chip (M4 Max 128GB) and I've been trying to optimize a completely purpose-built implementation with Fable and this is just not possible. Even if you could reach the full 576GB/s, it's just physically impossible to exceed these numbers with the model's architecture: 2 bpw - ~85.7t/s 3 bpw - ~58.0t/s 4 bpw - ~43.9t/s 6 bpw - ~29.5t/s 8 bpw - ~22.2t/s 16 bpw - ~11.2t/s without cheating. You'd have to exclude layers, skip operations, etc. basically do stuff the model wasn't trained for. And speed collapses so fast with context that even 2 bpw would be looking at ~37.6t/s after just 128K tokens. MTP only improves the situation by up to 2x in the ideal case, while drastically reducing the performance floor. While optimizing a 9B model on this hardware, I've found that the GPU just doesn't have enough FLOPS to handle speculating more than one or two tokens ahead on a single stream, regardless of quant level, simply because of the arithmetic cost of the forward pass. The 27B model would be even more expensive than that, potentially such that it's already bottlenecked by the GPU itself rather than memory. I wouldn't get my hopes up for the 35B-A3B either. Not only is it reportedly much less intelligent, but I hit a similar ~85t/s wall in practice (again with highly specialized inference). Without speculation I can reach around 120t/s on Qwen3.5-9B and with n-gram speculation (not even MTP; this derivative didn't come with one) around about 150t/s on average. This is on the very very edge of what I'd consider acceptable for me to even consider using such a compact model. YMMV due to the silicon lottery but the situation isn't good.
- minimaltom 2mo agoWhat is bpw? Also whats your cutoff for 'acceptable' speed? I would have said 25tok/s.
- ThouYS 2mo agoI am so happy right now, qwen3.6-27b was an absolute game changer. To see another one in the same league.. phew
- jedbrooke 2mo agoI hope the bonsai team makes another 1bit quant of this model (or releases code/instructions on how to do it), using the Qwen3.6 27B on my 16GB mac mini has been wild . The 1bit quant feels like opus level… for the first couple turns. Then it has trouble eg switching from plan mode to act mode. This is mostly mitigated by starting a new session. (tbf this limitation is called out on the hf page) I saw unsloth has 1bit quants too so I might check that out, anybody have experience with those?
- spwa4 2mo agoSounds like you need to check what the max context is set to ...
- jedbrooke 2mo ago100k is all the context I have ram for, this is with any auto-compact turned off. This is using Cline in vs code. I’m sure I could tune the system prompt and mode switching more to work better with this specific model, but I haven’t gone down the custom harness rabbit hole yet. And this is also specifically for the 1bit quant version. I don’t think the fp8 or even fp4 versions have this issue, but I haven’t tried those much
- prometheus1992 2mo ago16GB mac mini - what chip? m4 pro i assume?
- jedbrooke 2mo agonope just my normal m2 mac mini. I bought it just as a normal computer to do my taxes and whatever, so it’s mind blowing that I can run this kind of AI workload on it. Well, “run” might be generous, it gets like 3-5tok/s I’m working on a setup that’s more geared towards running tasks overnight so the slow tok/s doesn’t matter as much
- prometheus1992 1mo ago
- LeBit 2mo agohttps://xcancel.com/Alibaba_Qwen/status/2088280182356611304 https://xcancel.com/Alibaba_Qwen/status/2088280182356611304
- looksjjhg 2mo agoI could kiss you right now
- newaccount670 2mo ago[dead]
- brcmthrowaway 2mo agoThis with ddg mcp to fill in world knowledge. Are local models the future when computer architectures catch up?
- pu_pe 2mo agoSeems to be SOTA for its size. Hopefully independent benchmarks will come soon.
- ramon156 2mo agoneed another fable uncensored merge with 3.8, really curious what it can deliver
- deleted 2mo ago[deleted]
- xlayn 2mo agoThe file "Just loads" on llama.cpp, the Unsloth https://huggingface.co/unsloth/Qwen3.8-27B-GGUF https://huggingface.co/unsloth/Qwen3.8-27B-GGUF is an MTP file, I see mostly the same speed on pp and generation. There has to be something wrong with those benchmarks, I find extremely hard to believe a 27B model can work similar or exceed opus 4.6.
- minimaltom 2mo agoWorth distinguishing knowledge/task benchmarks from IF / agentic. It doesn't seem out of the question that you can have a small model thats generally good at instruction following and long-horizon agentic, as usually in those cases any requisite knowledge is in the context. Most of the benchmark improvements afaict are in agentic and instruction following benchmarks.
- cyanydeez 2mo agoI think you've been drinking the "LLMs only improve by adding parameter counts" that SOTA labs are selling VCs to build data centers so they can keep eating through cash to their own benefits. To the countrary, the reason Chinese models are excelling in the smaller area is because there's tons of fat in closed source models because of the crazy cash being thrown around. There absolutely is space to improve intelligence and capabilities without lathering on more and more parameters.
- selectively 2mo ago[dead]
- ThouYS 2mo ago3.6-27B on little-coder was already mind blowing. looking forward to this guy!
- mickeyp 2mo agoModel benchmarks are useful, to a point, but it is the long tail of things you do with the model that determines if it's good at a wide range of activities. Ant/OAI, to their credit, build their models -- even the small ones -- so they follow instructions and do tool calling well, without the system prompts confusing them. This is especially important for long-horizon tool calling. So one open weight model might "meet" Opus or whatever on benchmarks, but then fail to follow a simple answer format and also tool call correctly. The models are whipped to within an inch of their lives to strictly adhere to their post training quality gates.
- yassa9 2mo agoCan anyone who has that specific personal test he tries on different models , and tries this model , to tell us here if possible , how good or bad is this new model ? compared to others ? I only trust those users genuine personal tests
- alyandon 2mo agoThere is a down to earth guy on YT that performs a series of tests against LLMs running on non-god-tier commodity hardware. He will likely be testing this soon enough. https://www.youtube.com/@lukesdevlab https://www.youtube.com/@lukesdevlab I don't know if that is what you are looking for or not and as always your experiences may be different.
- yassa9 2mo agothaaanks man, this channel seems really informative, although < 10K subs only !
- alyandon 2mo agoIt's a relatively new channel - but yeah - I feel the guy puts a lot of effort into what he does and deserves more subs.
- xscott 2mo agoSo much potential for that channel. He's got a nice range of tests and a no nonsense presentation style. However, watching tests of heavily quantized models that weren't designed for it (non-QAT) is frustrating. There's no way to tell if the actual model fails because it's dumb or if the lobotomy made it that way.
- alyandon 2mo agoI noticed he does pay attention to feedback on his videos and I think some people have pointed that out.
- 1mo ago
- Mr_Eri_Atlov 2mo agoThis is the homelab model hands down
- synergy20 2mo agoI wish this can run directly on my RTX 4090, seems like 30B is the sweet spot for dense model to run locally, sadly RTX 5090 is very expensive and I need a new PC and new power supply(and UPS) to run that, adding a second RTX 4090 is another option, but not sure if my PC can do that yet.
- piyh 2mo agoQwen 3.6 is ~$2/m tok, 3.8 should be drop in replacement. Gemma 31B is $0.34/m tok. The price differential on these models is massive on openrouter.
- jjice 2mo agoWhere do you see that? From what I can see on Open Router, Qwen 3.6 27B (the closest dense equivalent to Gemma 31) is $0.28/m. Am I missing something? https://openrouter.ai/qwen/qwen3.6-27b https://openrouter.ai/qwen/qwen3.6-27b
- deleted 2mo ago[deleted]
- satvikpendem 2mo agoThey're comparing Qwen 3.8 Max to Gemma 31B, fundamental mistake.
- deleted 1mo ago[deleted]
- piyh 1mo agoYour own link says: In / Out Price $0.289 / $2.40per 1M
- SparkyMcUnicorn 2mo agoYeah, I would appreciate if someone could make sense of the pricing differences between these models. How can a provider run DSv4F at lower cost than a 27B dense or 35B A3B model? Does it come down to utilization and/or specific model tricks and efficiencies (attention, kv cache, etc.)? DeepInfra prices: Qwen 3.6 27B: $0.32 in / $3.20 out Gemma 3 27B: $0.08 in / $0.16 out DeepSeek V4 Flash 0731: $0.08 in / $0.18 out Qwen 3.6 35B A3B: $0.10 in / $0.95 out https://openrouter.ai/qwen/qwen3.6-27b https://openrouter.ai/qwen/qwen3.6-27b https://openrouter.ai/google/gemma-3-27b-it https://openrouter.ai/google/gemma-3-27b-it https://openrouter.ai/qwen/qwen3.6-35b-a3b https://openrouter.ai/qwen/qwen3.6-35b-a3b https://openrouter.ai/deepseek/deepseek-v4-flash-0731 https://openrouter.ai/deepseek/deepseek-v4-flash-0731
- theanonymousone 2mo agoI'm wondering whether any provider can offer this for cheaper $/token than the new DSv4 Flash, which is both cheaper and smarter :/ Completely local use is a different story, of course.
- ramon156 2mo agoPeople will claim it's not comparable to Opus despite it beating the score. I'm not sure I disagree, but I'm also unsure whether I care. Most new models nowadays are "good enough". I cannot complain because I'd rather spend that time improving my prompts and docs. Opus might be a _slight bit better_ at picking up vague hints, but it's also extremely expensive, and I hit the 5 hour limit way too quick. I care a lot about speed and efficiency right now. For my setup I would like to have 2-3 different model families. I've settled on GLM-5.3 (formerly Deepseek v4 pro 0813) for architecting, Deepseek V4 Pro 0813 for developing, and Gemini flash lite (any recent cheap model) for repo scouting. I'll add another one in the mix for reviewing (in this case Gemini 3.7) and that's all I need. I've tried most models except Grok. Qwen is too expensive IMO (Alibaba Cloud subscriptions are hard to come by and I'm not spending 50 euros a month for a tool, so 18 euros it is). If it ever becomes efficient enough to run locally I will definitely look back. Claude is slow and expensive (the cache hit prices are absurd). OAI is pretty good, I might add it to my arsenal seeing how cheap it is. These opinions change every day. Last week I would've never picked Deepseek until I read about the pricing. even post aug 16 it's worth it (although it's getting close to gemini pricing). Right now my costs are 12 euros a month (z.ai) + whatever deepseek consumes. This typically isn't more than 8 euros a week. 44 euros a month and I have a setup that is doing pretty well.
- hypfer 2mo ago> I've settled on GLM-5.3 (formerly Deepseek v4 pro 0813) for architecting Dude, GLM-5.3 released _today_. The phrasing "I've settled on" is incorrect for this context.
- ramon156 2mo agohence the "former deepseek v4 pro". I tried it out this morning and have had no complaints. I already liked glm 5.2
- hypfer 2mo agoThe sentence still doesn't make sense, because "settled on" implies a long testing phase with a verdict eventually emerging out of that. What you're currently doing is "testing out"
- filup 2mo agohttps://news.ycombinator.com/item?id=48403639 https://news.ycombinator.com/item?id=48403639 my prediction was way too far out. 4.6 at home! Woo.
- T0mSIlver 2mo agoUnsloth Q4_K_M on a single 3090, llama.cpp "Generate an SVG of a pelican riding a bicycle" first try https://www.reddit.com/r/LocalLLaMA/comments/1voa3ch/comment/p3o0om9/ https://www.reddit.com/r/LocalLLaMA/comments/1voa3ch/comment...
- btbuildem 2mo agoMost people cannot draw a bicycle that well!
- jlkivey 2mo agoNote: on the model card the comparison to Opus is Opus 4.6 Max, not 4.7
- RobertasTa 2mo ago[flagged]
- Casteil 2mo agoOne thing a lot of people don't seem to factor when hyping Qwen is how much models like this tend to 'overthink' with seemingly endless 'second guessing'. 3.8 seems no different from what I've tried thus far. As capable as it is, it's hard to justify using it when a competing model (e.g. Gemma4:26b-a3b) can consistently achieve the same or similar response with only 1/10th as many 'thinking' tokens, achieve much higher tokens/second, and take a small fraction of the time. I suppose 'YMMV' depending on your use case. Also, I haven't used it enough yet to see if it's prone to infinite looping, but its predecessors sure were.
- lrvick 2mo agoUse 3.6 27b as a daily driver for months with charmbracelet crush. Gemma 26b-A3b is not even remotely comparable in terms of coding for me. YMMV depending on how you work, what harness you use, etc I suppose.
- NamlchakKhandro 2mo agoreally? crush.... it's trash harness compared to pi. I didn't realise there are people out there unironically using crush
- lrvick 2mo agoYeah. I use it to do extensive work on full source bootstrapping, deterministic operating systems, compiler debugging, kernel debugging... all with one tiny go binary without endless NPM deps like pi (which /I/ regard as trash) Works better than opencode (pi based) or anything else I have tried for my needs, and by far the prettiest and easiest to reason about what is going on. But I will bite. What does pi do today better than crush for your use cases?
- Rayosay 2mo agoWhat's wrong with it? I like Crush. It's hard to find good harnesses that don't pull in mounds of Javascript like Pi and OpenCode.
- 2mo ago
- minimaltom 2mo agoArchitecture thread! Afaict they continue to use gated attention + delta net, which was also adopted+adapted by K3, but im surprised theres no improvements to the residual stream (deepseek are using manifold hyper-connections, kimi have attention residuals) ? Perf improvements seem to all come from training?
- anana_ 2mo agoAs was the case with GLM 5.3, it seems that there is still much juice to be squeezed from post-training
- fintuner 2mo ago[flagged]
- irthomasthomas 2mo agoWhy don't qwen/alibaba host the model themselves? I was looking forward to trying it on their coding plan. Google are the same way with their Gemma models.
- spwa4 2mo agoPretty sure you can use Gemma models on Google's "Vertex AI".
- arjie 2mo agoI use the Qwens as a vision model for my DeepSeek V4 Flashes to handle. But the Qwens run on old RTX A6000 Ampere. Does anyone know if there's any news about INT4/AWQ quants for the RTX A6000?
- ericd 2mo agoWas recently thinking about doing something similar, do you basically just have the qwens describe what they see for the flashes? Was considering adding a LoRa/vision head to Flash, but seems like it could take a while to get it right. If DSv4 Flash was multimodal, I’d probably be done model shopping for a while
- arjie 2mo agoSame, with a multimodal DSv4 Flash I would just stop paying attention to things. Very smart, and at 260 tok/s it's too fast to care about anything else. If you ever graft something like that I would love to hear about it. Yes, I have a very dumb flow. The harness has a describe_image tool that takes an image and a prompt and so DSv4 Flash uses it to get an idea of what it's looking at.
- ericd 2mo agoYeah, I might just replicate what you're doing. Main issue right now is just finding spare vram to actually run another model in parallel... And yeah, if I train up a vision adapter somehow, I'll try to put it up/post about it, seems like we're getting the killer apps for local LLMs right now, where it's just feasible enough if you're enthusiastic enough to be a bit economically irrational, and just useful enough to sort of rationalize.
- bertili 2mo agoWow. Speed improved as well. 200t/s on a RTX 5090! https://x.com/sgl_project/status/2088281320422322413 https://x.com/sgl_project/status/2088281320422322413
- hypfer 2mo agoSince it might be helpful to some, here's my current commandline for llama.cpp running on an RTX 4090 with my monitor moved to the iGPU to free up all of its VRAM. llama-server -m Qwen3.8-27B-IQ4_NL.gguf --mmproj mmproj-BF16.gguf -c 170000 --parallel 1 -ngl -1 --cache-type-k q8_0 --cache-type-v q8_0 -b 1024 -ub 512 --flash-attn on --no-context-shift --no-mmproj-offload --spec-type draft-mtp --spec-draft-n-max 5 --spec-default --cache-type-k-draft q4_0 --cache-type-v-draft q4_0 --threads 24 --jinja --reasoning on -fit off Identical to the qwen3.6 config. With a prompt like "svg owl" (which can reuse quite a lot compared with creative writing or similar, so ngram-mod shines), I get about 70-80t/s like this, with a memory overclock of about 1.5GHz
- reilly3000 2mo agoThanks for posting! Have you had any success with running without kv cache quantization? Is there a noticeable difference in quality without any? I would assume that would eat into context but 170k is pretty generous!
- hypfer 2mo agoAccording to this shitty vibecoded thing "I" built https://hypfer.github.io/will-it-fit-llama-cpp/ https://hypfer.github.io/will-it-fit-llama-cpp/ (and I guess according to math too), FP16 K/V would give me something like 90k context at the same model quant, which doesn't really fit my usage. But maybe someone else has experience to share there
- nubg 2mo agojust to clarify. yes YOU built it. just because you used some tool doesn't mean the idea, prompting, reprompting, babysitting was not your creative input and effort. put differently, if you put a random person infront of whatever model you used (say, a 50yo receptionist at a pharmacy in india), they would not have been able to create that, because they would have lacked the motivation, idea, background knowledge, taste, etc to create such a thing.
- 2mo ago
- mraza007 2mo agoMan what a week, We just had GLM 5.3 that came out and then we had smaller local model Qwen3.8-27B from Qwen Just tried using Pi Agent and looks very promising
- cmrdporcupine 2mo agoI found this kind of amusing while running it (using Pi as the harness). Don't know if this is evidence of intense fine tuning from Claude but it smells like it... " The user wants me to explore the repository at XXXX and report back. Let me start by understanding the project structure, reading the CLAUDE.md file, and getting a general overview of what this repository is. Let me start by reading the main project documentation and exploring the directory structure. I'll take a look around this repo. Let me start by getting a lay of the land. read resource CLAUDE.md (ctrl+o to expand) ENOENT: no such file or directory, access 'XXXX/CLAUDE.md'"
- deleted 2mo ago[deleted]
- steffi_oliver 2mo ago[flagged]
- Almondsetat 2mo agoThe $1500 Intel B70 with 32GB of VRAM can run this model at max context with good performance, btw. If you don't want to drop $5-10k for running DeepSeek this is your best budget option for local refactor/small scale dev help
- bogzz 2mo agoOh, can it work with the /v1/completions/ auto-complete endpoint?
- Almondsetat 2mo agoSorry, I wrote autocompletion by force of habit. I simply meant it can complete code you have already created a structure for, which personally is very nice
- bogzz 2mo agoI thought so, but thanks for the clarification. I am a little bit disappointed that local autocompletion models have been left by the wayside in favor of models post-trained for agentic coding. Both Codestral and Qwen-2.5-coder are more than a year old at this point, but local auto-complete seems to me to be such a great usecase.
- gered 2mo agoThe latest Qwen models (including 3.8 27B) do still support FIM-style in-editor code auto-completion if that's what you're looking for. I wouldn't want to use a large dense model like 27B for such a task (since FIM-style auto-completion really works best with low-latency responses), but it works.
- LeBit 2mo agoI understand the B70 is a bargain vs AMD and especially nVidia offerings, but to me it feels like I would be buying something that would feel too limited in less than a year. 48G would be much more confortable. And I know the 96G nVidia cards are selling for over 10k$. The future can’t arrive fast enough!
- tristor 2mo agoI'm hoping to see folks distill this with current generation Opus / Fable reasoning traces. I have had my best results locally so far from Qwopus (Qwen 3.6-27B w/ Opus 4.6 reasoning distilled). This looks GREAT and I am definitely setting this up later today.
- satvikpendem 2mo agoAs usual, the Jinja templates are messed up so use this [0] to reduce or turn off thinking, fix tool calling, keep a 100% KV cache hit rate, etc. [0] https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
- skrebbel 2mo agoI'd love to understand this more. Are you saying the Qwen team spends their very impressive human and compute resources on publishing these amazing models and then botches the chat template with mundane bugs? Like maybe I just misunderstand what's the hard part but wouldn't you assume that people who can put together an impressive model can also write a proper jinja chat template for it?
- Der_Einzige 2mo agoYes yes, oh god yes. They also spread FUD in the form of terrible recommended sampler settings. If you're using llamacpp, turn on top-n-sigma with sigma of 1, turn off top-p/top-k. You'll thank me later.
- skrebbel 2mo agoHow is terrible settings a case of FUD?
- MrDrMcCoy 2mo agoFor those us us who don't know, what do those parameters do and why are they better?
- deleted 2mo ago[deleted]
- suprjami 2mo agoTemperature, top-up, top-k, min-p all control which token the model predicts next and how likely it is to select one token over the other. You might understand this as "The capital of France is..." and the model isn't always going to select "Paris". Sometimes it will start a descriptive sentence or even get the answer wrong. That selection of the next token is what these settings control, and lots of sub-optimal selections compound over time to produce a junk response.
- naasking 2mo agoCan anyone confirm whether this new Qwen release is any more concise when thinking? Overthinking was the biggest (only?) downside of the Qwen models.
- lossolo 2mo agoWhy weren't the points merged again from the "dupe" thread that had 289 points? https://news.ycombinator.com/item?id=49299684 https://news.ycombinator.com/item?id=49299684 What a weird mechanism. If someone is judging a thread/topic/event impact by the number of points it got, then doing this unfairly degrades that thread. It should have deduped by user and combined the 168(at the time of writing this comment) + 289 points. Just add the twitter link from the previous thread as an additional link in the description, like you normally do, move all the points over, and remove the old thread.
- davidw 2mo agoI don't know much about the production of these models. How hard would it be to 'fork' something like this and have it not be full of CCP indoctrination?
- regularfry 2mo agoLook for `heretic` fine-tunes in the next couple of days.
- syntaxing 2mo agoWould I be surprised there’s bench maxing happening? Yes. But some users also use Q4 quantized and complain how dumb local models are.
- literoldolphin 2mo agoWhy is anyone even using video cards these days? You may as well be burning cash. This is the perfect candidate for just splattering it on your nvme and then reading it off there and into memory. All of these run perfectly fine on simple m4 silicone: https://github.com/drumih/turbo-fieldfare https://github.com/drumih/turbo-fieldfare https://github.com/leonickson1/Swiftlet https://github.com/leonickson1/Swiftlet https://github.com/sqliteai/warp https://github.com/sqliteai/warp
- awkwardpotato 2mo agoThose are all for MoE models. And I prefer measuring my tokens in t/s instead of s/t
- ferrouswheel 2mo agoLol, "burn money on apple hardware instead!"
- literoldolphin 2mo agoAnd yet it's also a laptop you can basically take anywhere unlike a giant video card with 1000 watt power supply requirements.
- bigyabai 2mo agoNvidia and AMD both have their own unified memory laptop SOCs, now. Apple Silicon's GPU is relatively weak, it's one of the less-efficient ways to use 100w for compute. Even the fastest Apple Silicon chips like the M5 Max and the M3 Ultra still put up worse GPU compute performance than last-gen laptop RTX 4080 chips. And they don't scale, the largest M3 Ultra cluster you can configure is still ~2,000x smaller than a DGX SuperPOD. There's a reason Apple discontinued their rackmount hardware, there's very little demand for Apple Silicon in the datacenter.
- literoldolphin 2mo agoAs I mentioned above you can't take the data center into the cafe somewhere. We're talking about running local models here.
- monkmartinez 2mo agoQwen3.6-27B has been the main LLM powering my little agentic stack. I have adopted the test and verify approach to any models allowed to run on my machine. When the "heretic" version drops, I will fire up the harness and test. Super excited to see how it stacks up against Qwen3.6!!!
- deleted 2mo ago[deleted]
- Balinares 2mo agoI wonder if Anthropic and OpenAI possibly missed the window to go public. A 27B open-weight model trading blows with the SOTA from just half a year ago is not great news for trillion-dollar investments...
- brcmthrowaway 2mo agoThis is why they've been making bank on the secondary market. They can retire now.
- dannyw 2mo agoTrading blows in some benchmarks is a bit exaggerated. If you try the model, `xhigh` is basically feels like the `max` mode (i.e. massive thinker and extremely presistent), and the amount of world knowledge and intent understanding is nowhere close to an Opus class model even from 6mo ago. It's still very useful, and it'll probably displace a good bit of API spend; but it's not really "trading blows with SOTA from just a half year ago". A bit overblown on the Anthropic/OpenAI has missed the window I reckon. Also, on the open weight frontier side, Kimi K3 is pretty expensive, and Deepseek V4 Pro/Flash is getting a little less juicy with price increases.
- jdgoesmarching 2mo agoLacking world knowledge is fine for me, I rarely want to rely on the model’s training anyway when there are plenty of great search options to integrate with.
- simjnd 2mo agoWhy would anyone rely on the world knowledge built into a model when the harness can just let it search for current information? Intent understanding is a big point for sure, but world knowledge I'm not sure I see a use case for it.
- Balinares 1mo agoThat's a reasonable point, so let me qualify: at a glance, it seems Qwen 3.8 27B can trade blows with Opus 4.6 on coding tasks, where trading blows doesn't necessarily mean it's a clear winner or even an equal, but does mean it'll at least hold its own and land a punch or two. (Which I still think is bonkers, FWIW.) Opus 4.6 is an especially interesting comparison point, I think, because it was a step change; IMO it's when LLMs became serviceable for coding. Yeah, pre-4.6 models did output code, and that code often superficially worked; and bringing it up to production standards still generally meant rewriting it entirely. Opus 4.6 is when that changed. From my early tests, it's looking like the public benchmarks are not misleading, and Qwen 3.8 somehow got there too, by and large. I've got a few personal tests. One is a mid-complexity one-shot, purposefully underspecified. Beyond a few minor bugs that it could easily fix once pointed out, Qwen 3.8 largely aced it. There are a bunch of things I'd improve, but that was true of Opus 4.6's output too, and by and large the code is clean and well structured. Also worth noting that I'm running Qwen 3.8 fairly aggressively quantized to fit in VRAM; I'd expect tighter results still from the full weights. Another test I ran is a variant of a common puzzle with an additional structural constraint that makes the usual solution inapplicable, so the model has to actively turn away from the well-known solution and construct a new one that takes the constraint into account. I've never seen a home model pass that test. Kimi K3 passes it, GLM 5.2 passes it (painstakingly). Qwen 3.8 struggles a lot... but does arrive at a correct solution. First time I see a home model do so. I haven't yet tested it on long multi-turn scenarios. In my experience, that's where pocket models are weakest against heavyweight ones, especially when quantized. That aside, it does seem like Qwen 3.8 can, in fact, trade blows with Opus 4.6. I don't know yet if it could replace it, and my money would be on no, but I may well be wrong about that considering how weirdly capable it is. Interestingly, Qwen 3.8's MTP layer is uncannily accurate too. It still gave me good results up to 6 to 8 predicted tokens, which boosts its speed so much it's competitive with Qwen 3.6 MoE. So that's another bizarrely impressive thing about it. And given all of the above, I do think that the trillions of dollars invested into OpenAI and Anthropic are becoming harder and harder to justify.
- kanemcgrath 2mo agoI think I am going to buy a second rtx 3060, as 27B has been just outside of my range for to long, and this looks like the parameter count tipping point
- apitman 2mo agoRunning it on 2x3060 now. Works pretty well but VRAM is tight. 4bit quants. 1x128k context, 8bit KV, MTP on.
- kanemcgrath 2mo agowhats the tok/s you get on that. I have heard a few claims of around 30-50 with mtp, but for how cheap the setup is I am surprised I don't hear more about 3060 stacks so I assume there has to be some catch.
- apitman 2mo agoI used GPT-5.6 Sol high to optimize it, and it claimed it was getting 50. I'm seeing ~40 on my goto smoketest: "Make me a vector add in CUDA". Funny side note. It successfully one shot the program, but it wasn't able to run it because there literally wasn't enough VRAM left to allocate CUDA memory. Watching it try to debug that was fascinating. I'm pretty sure it would have killed the llama-server (and thus itself) if it hadn't been running in a separate container.
- esotericsean 2mo agoNeed to upgrade to a second 3090! Slowly building up my local models with Krea2, MiniMax H3 (and their new Music3), and now Qwen 3.8
- btbuildem 2mo agoO joyous day!
- deleted 2mo ago[deleted]
- maherbeg 2mo agoDoes anyone have a https://tenstorrent.com/hardware/cards https://tenstorrent.com/hardware/cards to try it on?
- imagetic 2mo agoYes.
- gaigalas 2mo agoWaiting for the MTP version to pop up on Unsloth. Speculative decoding makes a huge difference. Been running quantized 3.6 at 110t/s on a cheap 5060Ti and quite happy with it. If 3.8 improves on it, it would be awesome.
- zazibar 2mo agoGood news, MTP support is already included in this release. Not sure why they haven't made this clearer.
- gaigalas 2mo agoI don't see an MTP entry on Unsloth though. Maybe it's not available in a lower quant I need for my poor GPU.
- greenicon 1mo agoIt's included in the model gguf itself.
- gaigalas 1mo agoMaybe I'm not seeing the speed improvements because it's a dense model, and I'm comparing it to the 3.6 MoE with MTP. Likely, I'm attributing the speed I see to something else. Honestly, if MTP is available, I don't know why it's so much slower here.
- deleted 2mo ago[deleted]
- dude3 2mo ago[flagged]
- rcarmo 2mo agoHmm. No MoE or active params weights means this will run _slow_
- fr2029 2mo agoWill there be an A4B MoE?
- natch 2mo agoApart from model performance, what harness are people using to come close to Claude Code or Codex workflow styles with tool use, conversations, loops, remote control, etc.?
- yalok 2mo agoand more specifically - what harness is known to be the best fit for Qwen local models, and are there any evals/benchmarks for harness+model pairs?
- chillaranand 2mo ago"Generate an SVG of a pelican riding a bicycle" - generated a promising image at first shot. https://avilpage.com/qwen-3.8-27b.html https://avilpage.com/qwen-3.8-27b.html
- swalsh 2mo agoWOW, my first try running on my 2 3090's, it was a bit slow... but it FEELS like opus 4.5, i gave it an image and a broad overview of what I wanted it to build, and it built the whole thing from beginning to end.
- venusenvy47 2mo agoFor your setup, do you have both 3090's in parallel for the inference of the model?
- swalsh 2mo agoYes they run in parallel via LMStudio (250k context)
- XCSme 2mo agoWhy slow? I see ~50tps on a single 3090
- jhonof 2mo agoYeah this is the first model I have been able to run locally that actually feels useful, this is unreal I am considering cancelling my claude sub and going to just api (maybe GLM?) for really hard tasks.
- apitman 2mo agoCheck out OpenCode Go as well. They give some Kimi K3, Qwen3.8 Max, and GLM5.2 (probably 5.3 soon?) usage which may cover your needs for $10/mo
- jhonof 2mo agoYeah I was thinking open router but I will look around at options, I genuinely think this model is good enough for like 90+% of my use cases, and the top frontier models are still not that good at architecture so I have to do that myself still so I won't be losing out.
- g023 2mo agoAll this performance at such small model sizes, why are the API fees so high for the AI monopolists on this side of the world?
- jjcm 2mo agoImage->html test for this. Original images: https://image.non.io/neonRamenDesigns.webp https://image.non.io/neonRamenDesigns.webp Qwen 3.8 build: https://html.non.io/neonRamenQwen3.8-27b https://html.non.io/neonRamenQwen3.8-27b Overall I'm very impressed with how well this did. It's a big improvement over 3.6, and it feels on-par with some much, much larger models. I think this one is on-par with Gemini 3.7 Flash. One thing to note - the build for this on my RTX 6000 pro blackwell took a long time. Easily one of the longest builds I've done. It took around 2 hours to build the site. Obviously we'll have some quants for this soon that will accelerate things, but I was still surprised with how long it took. Comparison builds from this week: https://html.non.io/neonRamenGemini3.7 https://html.non.io/neonRamenGemini3.7 https://html.non.io/neonRamenGLM5.3 https://html.non.io/neonRamenGLM5.3 (note: non-multimodal)
- tomr75 2mo agowhat token/s?
- jjcm 2mo ago27 t/s. I suspect there will be significant speed ups in the coming weeks.
- apitman 2mo agoYou should be getting way more than that on a 6000 pro even today. I'm getting 40tok/s on a pair of 3060s. You can ask a SOTA model to optimize your setup for you.
- nojs 2mo agoAny idea why it’s so slow? the entire model should fit in the vram of one card.
- graceful6800 1mo agoThe NVFP4 quant is completely broken, so I'm not shocked that other quants aren't fully there yet. Give it some time to cook.
- dofm 2mo agoThere's a real change (compared to 3.6) in the way it writes in thinking — it drops words like "to" and "we" in "We need to", talks generally in note form, drops the/and all over the place, avoids "for". "Need be helpful concise", "Need maybe not overdo", "Need ask!" Almost caveman. I have an (unsourced, vague) suspicion that this rather unique thinking trace pattern is actually hobbling the MTP predictions, which seem to perform poorly. Other notes: it uses the trick of repeating the prompt in the thinking trace. It also worries about hidden chain of thought appearing in the final answer. It talks about "desired oververbosity 9", which is new. A bit GPT-ish. It is being extraordinarily thorough in thinking through one of my code requests, but I don't know if the net result will be any better than the 35B MoE. I asked it to ask me clarifying questions — it did, and it offered me a list of defaults I could simply agree to. I don't think it is necessarily overthinking in the looping sense, but it is in the being exhaustive sense. I need to explore how it does with a tighter reasoning budget. I am impressed but I am definitely in Camp Please-35B-A3B-When? here, because on an M1 Max this isn't really practical. I hope they do one, though I think they may not.
- dofm 2mo agoAnnnnd the code of my WP code test is not better. It is bushy, overcomplicated, and has gone around the houses to do stuff it would not need to do if it hadn't overthought. Oh dear. I need to try to understand what is going on here.
- eek2121 2mo agothinking is set to max by default. I bet that turning it down would solve this.
- dofm 2mo agoI think so too — it is something to test, for sure. ETA: a bit of testing before I climb the wooden hill to Bedfordshire. LM Studio doesn't seem to display the little dropdown to set reasoning effort, so I bodged the chat template on load to get it to choose 'medium'. As soon as you switch away from xhigh, it goes back to thinking in normal sentences like Qwen 3.6, rather than in sort of quasi caveman. And you get all the Wait, Actually, No wait… stuff back. And it is behaving a lot more like it used to. So that is pretty interesting.
- sheepscreek 2mo agoBetter than Opus 4.6 at computer use? Comparable with it for SWE? Am I reading this right? I’ve heard rumours about AI shops optimizing for benchmarks. I also don’t think Qwen/Alibaba would be crazy enough to claim something unless there is some truth in it. Would love to see a side-by-side with Opus 4.6 on categories where Qwen 3.8 27B aces it.
- simonw 2mo agoAbsolutely the best pelican I've seen from a model that runs on my laptop: https://tools.simonwillison.net/markdown-svg-renderer#url=https%3A%2F%2Fgist.github.com%2Fsimonw%2Ffc909bea4fecf752c7bf9bad0e9dbf2a https://tools.simonwillison.net/markdown-svg-renderer#url=ht... Bicycle is the right shape. Pelican beak is excellent. Nice background. Most importantly, the pelican has one leg on each side of the bicycle - that's very rare. (No chain on this bicycle though - in the reasoning trace it says "already chainstay... skip chain detail; maybe a small chainring.") I ran that on an M5 Max MacBook Pro using LM Studio and their 17GB GGUF: https://lmstudio.ai/models/qwen3.8 https://lmstudio.ai/models/qwen3.8 It took 21 minutes(!) and used 22,276 reasoning tokens to produce 3,223 tokens of output. (For the "they're training on your benchmark now" crowd, all of that cheating didn't prevent it from spending 20 minutes thinking about the task first! You can see the reasoning trace in the link I shared.) For comparison, here's one I got from qwen3.8-2.4t-a95b on OpenRouter, which is pleasingly animated: https://tools.simonwillison.net/markdown-svg-renderer#url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F557016f0895b2abb4b9957caec781734 https://tools.simonwillison.net/markdown-svg-renderer#url=ht...
- simonw 2mo agoFor anyone who followed yesterday's Gemini 3.7 Flash pelican which rendered in Safari but not in Firefox or Chrome... https://news.ycombinator.com/item?id=49289112#49290012 https://news.ycombinator.com/item?id=49289112#49290012. - that turned out to be my fault, not the model. My SVG rendering software was stripping some attributes. Here's the Gemini 3.7 Flash pelicans in the fixed renderer: https://tools.simonwillison.net/markdown-svg-renderer.html#url=https%3A%2F%2Fgist.github.com%2Fsimonw%2F6779a22d5e7bb6bdf29936f1600a5259 https://tools.simonwillison.net/markdown-svg-renderer.html#u...
- Einenlum 2mo agoDamn That's very good
- theplumber 2mo agoGemini is worse
- 2mo ago
- cloudengineer94 2mo agoBeen trying out Qwen3.8-27B-Q5_K_S_20GB and it's quite interesting it's behaving very well. Going to give it some coding tasks and see how it goes. We been eating good at LocalLlama this week.
- seanmcdirmid 2mo agoI'm struggling to figure out what to use this for. From the intelligence benchmarks in OMLX. If only they would release another MoE model. Intelligence Benchmark Comparison --- Detail --- Model: scottlowry--Qwen3.8-27B-oQ4e-mtp Benchmark Accuracy Correct Total Time(s) Think -------------------------------------------------------------- GSM8K 93.3% 28 30 282 No MATHQA 46.7% 14 30 26.3 No HUMANEVAL 96.7% 29 30 156.5 No MBPP 83.3% 25 30 71.5 No LIVECODEBENCH 43.3% 13 30 1040.4 No Model: stamsam--Qwen3.6-35B-A3B-Claude-4.7-Opus-Reasoning-Distilled-MLX-oQ4-MTP Benchmark Accuracy Correct Total Time(s) Think -------------------------------------------------------------- GSM8K 96.7% 29 30 51.9 No MATHQA 60.0% 18 30 9.1 No HUMANEVAL 83.3% 25 30 82.9 No MBPP 80.0% 24 30 29.6 No LIVECODEBENCH 36.7% 11 30 283.7 No
- walrus01 2mo agoTry manually asking both more discrete esoteric knowledge questions. Or use benchmarks which are less coding focused. The 3.6-35B-A3B with post-training may do well in coding type benchmarks and math but the density of its knowledge falls off in my experience (vs 3.6 27B dense Q8-K-XL unsloth GGUF) when you need to use it for less commonly used domains of knowledge.
- seanmcdirmid 2mo agoMy use cases try to avoid accessing world knowledge in the model (I give it access to web search for some adhoc RAG), and ya, I'm just focused on coding so that's the only place I'm looking at right now.
- walrus01 2mo agoI think you may find that the dense 27B also does better if challenged with more rare coding tasks, less common or weird languages or things that aren't well represented in the active 3B parameters of the MoE model (eg: NOT css, javascript, python, c++, etc).
- z_rho_one 2mo agoBeating or comparable to Opus 4.6 in benchmarks. Opus 4.6 was released in February, 2026. So if we still want to talk about a "6 month difference" between Chinese and American AI, the sentence should now be: Chinese (small model) AI is 6 months behind American (largest model) AI.
- jared0x90 2mo agoGiven that glimmer only caught up-ish to 3.6 how far behind is American (largest model) to American (largest model) ?
- jacquesm 2mo agoThe default reasoning is set to 'xhigh', if you want to compare with the past or reduce the time (if you can take the hit in output quality) then you can pick 'high' or 'medium' as well.
- spijdar 2mo agoI wonder how this practically compares with Muse Glimmer, especially quantized. I've got an RX 7900 XT (20GB of VRAM) and I can run glimmer with a full 128k context window with the draft model at 65-80 tok/s. This model, on the other hand, I get about 30 tok/s with a 30k context. Raising the context or loading the draft layers for MTP drops performance to 9-15 tok/s. So I wonder how big the "real world" delta between Glimmer and Qwen is here. I can already run 3-bit DSv4-flash at 9-15 tok/s with 100k~ context, and I suspect it would outperform 4-bit Qwen 3.8 27B here. I'll have to experiment and see if I just made a stupid mistake somewhere, but it looks like Glimmer might make more sense for the comically specific niche of "20GB VRAM".
- jakswa 2mo agoI'm in the exact same boat with a 7900 XT and a good Glimmer 30B experience. I was really hoping qwen 3.8 would bring some memory/space efficiency savings along the lines of whatever is going on with Glimmer 30B. I have been surprised that a 30 billion model fits and runs better (at higher unsloth quantization! UD-Q4_K_XL fits!) than a 27 billion model.
- harhargange 2mo agoI too purchased the 7900xt as it was cheap with a lot of vram. Qwen 3.6 27b gives me 30 tok/s
- marius_ 1mo agoFrom my limited tests, Qwen has better reasoning which is a bummer because Muse Glimmer is literally twice as fast.
- cloudengineer94 2mo agoHaving some really good fun with it so far with System Design (ERP) mainly SAP. I did notice if you go beyond Medium he starts overthinking like hell as per usual for a Qwen model.
- xlayn 2mo agoThe unsloth Q8kxl https://huggingface.co/unsloth/Qwen3.8-27B-GGUF https://huggingface.co/unsloth/Qwen3.8-27B-GGUF for some reason is looping and going crazy on the think part (I tried to search for an email to let the guys know but didn't find one)... I used the bartowski one and that one doesn't have that issue https://huggingface.co/bartowski/Qwen3.8-27B-GGUF?show_file_info=Qwen3.8-27B-Q8_0.gguf https://huggingface.co/bartowski/Qwen3.8-27B-GGUF?show_file_... that's using llama.cpp llama-server \ -m ~/somePath/Qwen3.8-27B-UD-Q8_K_XL.gguf \ -np 1 --kv-unified \ -fa on --no-cache-idle-slots --reasoning-preserve \ --temp 0.2 \ --spec-type draft-mtp,ngram-mod --spec-draft-n-max 3 --spec-draft-n-min 1 \ --cache-type-k f16 --cache-type-v f16 \ --chat-template-kwargs '{"preserve_thinking": true}' \ I tried playing with all the recommended parameters from the unsloth page with no luck... in one of the high fever ramblings it ended with amen... lol
- SwellJoe 2mo agoI'm not seeing that with that exact quantization from Unsloth, so far. I'm seeing a _lot_ of thinking before it starts doing, but it all seems pretty reasonable and not loopy (at least no more loopy than big models, with the expected "But, wait! I need to..." types of back-tracking). So, it's taking a long time, but I don't think it's doing anything pathological.
- dofm 2mo agoI would characterise it as obsessive, not loopy. It's definitely burning through a lot of tokens to ruminate about aspects of tasks that earlier models get done better seemingly on memory. I still need to understand that. Setting a reasoning limit does not seem to have good results, because it really seems to go down rabbit holes and that means that cutting reasoning off too early is going to punish the quality on anything it has not got round to pondering yet. But maybe I have to give it a bit more room. I have not tested in an agentic sense yet, just with my sort of pet queries in LM Studio, but it rather looks like it expects an agentic flow, because telling it that it's a helpful coding agent and changing the order of things in my prompts (telling it up front to ask any clarifying questions before detailing the rest of the prompt) has definitely kept its thinking a bit more on track.
- c7b 2mo agoFor those commenting on the long reasoning, it may be interesting to know that the reasoning effort is set to xhigh by default [0]. Other possible values are medium, low and none. Flag for changing it in llama.cpp below, but note that the long reasoning seems to contribute a great deal to the quality. --chat-template-kwargs '{"preserve_thinking":true,"reasoning_effort":"medium"}' [0] https://unsloth.ai/docs/models/qwen3.8#thinking--preserve-thinking https://unsloth.ai/docs/models/qwen3.8#thinking--preserve-th...
- 482937632992 2mo ago[flagged]
- brcmthrowaway 2mo agoIs oMLX or MTPLX supported?
- CMay 2mo agoCredit where it's due. Qwen 3.8 27B is only the second local model after Gemma 4 that managed to correctly reason through one of my private benchmarks. It took 5x as many tokens to do it and 12m30s with MTP enabled, but it did do it. Gemma 4 reasoned through it more implicitly, while Qwen 3.8 reasoned more explicitly. Laguna and Muse Glimmer failed hard on it, though they're useful for other tasks. The VRAM usage seems way less efficient than Gemma 4 or Glimmer though, with 32K of context taking 2.5GB of VRAM. With those, even with MTP or a DFlash model loaded, you could still fit 256k-768k of context. With Qwen 3.8 27B I can't even fit 128k if I quantize V to Q4_0. Maybe with some trial and error I can find some settings that perform well enough with a larger context window that it's still useful for longer tasks. Lots more testing to do, though I was getting some decent results out of Muse Glimmer which was more than twice as fast and supported huge context windows, managing to solve some bugs that Gemma 4 struggled with. I can't even begin to throw that task at Qwen, because just the prompt alone would use the entire context window and then it would reason for probably that same amount. If you've got a 32GB card, it should be a decent model even if it really is memory hungry. EDIT: Tried a few kv cache quantization settings, but it failed with those. I designed this benchmark to be pretty brutal in the face of KLD and any reasoning quality loss, so it's not too surprising. Gemma 4's QAT held up pretty well, at least and could consistently complete it.
- dofm 2mo agoI was quite impressed by Muse Glimmer, and while I am sure people will observe that it is less good on benchmarks, my first experiences with this new 27B have been somewhat exasperating, whereas testing Muse Glimmer was rather fun. I have not tested either in an agentic context, mind you.
- chr15m 2mo agoGlimmer is fun because it's fast, tight, and doesn't wander or waffle. My favourite local model so far.
- pram 2mo agoGlimmer works really well as an "explore" agent model (like in Opencode.) It seems to be extremely efficient at searching and collating that info, and executing commands. From my testing so far, Qwen 3.8 is better at code but it tends to meander and take forever if it has to look in a lot of places. Glimmer will use like ~1k tokens to formulate a plan and Qwen 3.8 will routinely go over 10k
- Valdior 2mo agoI am waiting for Qwen 3.8 MoE - last time 3.6 MoE was better on codding that just dense 3.6.
- simonw 2mo agoAnyone seen this show up in any APIs yet? I'd love to try it out faster than my Mac can run it.
- abidlabs 2mo agoThe Hugging Face page (https://huggingface.co/Qwen/Qwen3.8-27B https://huggingface.co/Qwen/Qwen3.8-27B) shows that Featherless supports the model (and includes the inference widget to try it out directly on the page)
- walrus01 2mo agoIs it just me or is 3.6 27B Q8 K XL (Unsloth) holding up better in sustained token/s rate as the context fill increases over time? The token/s rate seems to be much higher for a time period deeper into context than previously seen. At least as compared to 3.6 27B in the same quantization.
- crazyemeraldcod 2mo agoIts so smart!
- potus_kushner 2mo agohopefully for us mere mortals without $4k+ hardware a 35B MOE model will be released. or a new prism ternary bonsai model based on this one.
- Anonyneko 2mo agoIs there any way to turn off thinking if I'm using Ollama? In my particular case, the Ollama API (the software I want no-think for is tied to Ollama's bespoke API). If not, I'll stick to 3.6 for the time being...
- svdr 2mo agoWow. This model is so good, and we have GLM 5.3 (seems great voor security related work) and Deepseek. In a few months we'll have Fable/Sol-like capabilities that are not coming from the big US companies. I feel as a programmer that that is more than enough. How wil OpenAI and Anthropic survive when frontier model intelligence becomes commoditized?
- horacemorace 2mo agoThey’ll certainly try to stymie people by colluding with manufacturers until we get nvidia level hardware or LLM ASICs from the East.
- awb 2mo agoThe demand curve for speed and intelligence seems pretty steep to me. If you look at the hiring marketplace, being just marginally better than your peers can be very lucrative. If you’re competing on speed or capability as a company (or as an employee), you’re probably going to be willing to pay for the frontier.
- croes 2mo agoMost companies have a limited budget. Good enough with a lower price will win the masses
- HarHarVeryFunny 1mo ago> If you look at the hiring marketplace, being just marginally better than your peers can be very lucrative. I would say that in software this is completely false. Someone straight out of college, not very useful, makes 75-100K. Top level senior outside of FAANG is making twice that at best (and at least 10x more capable).
- Der_Einzige 1mo agoYou told on yourself about being either European or from a flyover state.
- kimsey0 2mo agoIf anyone else is running this on an RTX 5090, https://github.com/Neroued/ninfer https://github.com/Neroued/ninfer as inference engine gets me ~138 tokens/second, roughly double what I get with a naive llama.cpp setup.
- wincy 2mo agoAmazingly enough, Hacker News decided to show this to me as the top comment, I'm also running an rtx 5090 and trying it out now, thanks for the tip! Edit: Absolutely blazing fast! Getting 163 tokens/sec on WSL and it generated a pretty sweet Pelican. https://gist.github.com/hansale/ed9e73fe35165a58ea2af6b1632afdb5 https://gist.github.com/hansale/ed9e73fe35165a58ea2af6b1632a...
- sgt 1mo agoAmazing. I'll give it a shot on my 5090. I already tried using vLLM but it ran out of GPU memory. I guess it's likely Llama.cpp will work.
- deleted 2mo ago[deleted]
- searealist 2mo agoJust enable MTP on llama.cpp and you will get the same decode speeds.
- pulse7 2mo agoIt there anything similar for RTX 3090 and RTX 4090?
- mmlkrx 2mo agoI'm not sure about a 4090 but there is a fork for 3090s: https://github.com/Don-Chad/ninfer-3090 https://github.com/Don-Chad/ninfer-3090
- mirekrusin 2mo ago[dead]
- amazingamazing 2mo agoWith my 5080 laptop so close yet so far to using this stuff
- webbrain 2mo agowow
- webbrain 2mo agoit is amazing folks! it works like charm with 5090
- singingtoday 2mo agoPlayed with this on my Mac a bit today. Not bad!
- scgopireddy 2mo agoOn M5 Max I am only receiving 17 tok/sec, how do I maximize
- scgopireddy 2mo agoOn M5 Max, I am receiving only 17 tokens/sec. How do I maximize tok/sec?
- akg_67 2mo agoAre you using MLX version on M5 Max? qwen3.8:27b-mlx
- xcf_seetan 2mo agoHmmm, Just started using it and it is the first time a LLM tell me this: "You can write this yourself. It is not hard." lol
- RandyOrion 2mo agoThank you Qwen team for this release. Compared to closed weight (especially unreleased and access-limited) and open weight/source large sparse MoE LLMs/VLMs, open weight/source small dense models benefits public the most because they just reaches more people. Compared to Qwen 3.6, 3.8's thinking style changed drastically. With xhigh budget, it thinks a lot MORE, and longer thinking session directly translates to better performance. This tradeoff between performance and computation, memory, etc. is meaningful to me. However, because Qwen 3.6 and 3.8 share the same architecture, with 32GB vram, llama.cpp, IQ4_XS model, MTP and FP16 mmproj, I can only get 200k context, which is not good compared to 640k context of muse glimmer. Hopefully this problem will be solved in Qwen 4.0 release.
- scripthound841 2mo ago[dead]
- pdude444 2mo agoDo we really feel like it’s the governments job to regulate OPEN source AI. At rely health, we use OSS models in a HIPAA complaint and SOC 2 complaint environment to take advantage of asymptotically $0 intelligence to provide best in class care navigation . This should be industry standard -
- kofj 2mo ago[dead]
- rohan_tech24 2mo ago[dead]
- dexterlagan 2mo agoTested the model briefly with my usual eval: a couple of questions on general knowledge most small models often get wrong, then write a fully-featured todo list app in JS, then rewrite the same app in Rust with Tauri. Granted, most models are well trained on basic todo apps, but it gives me an idea of the basic SWE capabilities I can build on. As far as I'm concerned, if it can successfully setup a local git repo, write a todo list app skeleton that works, I can work with it. SWE: model is strong for its size. It one-shotted the Web app, had no bug. The Rust rewrite only had one bug (reordering didn't work immediately - fixed in one prompt). Committed locally then pushed to my GitHub (https://github.com/DexterLagan/RusTODO https://github.com/DexterLagan/RusTODO). Can't complain. If it can do that reliably, I can use it to make whatever I need. General knowledge: I usually ask 2 questions many small models get wrong: summarize Operation Trojan Horse by John Keel, give publication year. Summarize the Ariel school incident of 1994. Qwen3.8 got the first question right, along correct publication year, gave glorious detail on Keel's theory, but got the second wrong. It thought that school was located in the USA. Ah well. Performance on an admittedly overpowered laptop: 15 tokens/s in power save mode on this MacBook M5 Max 48GB, and 30 in performance mode. Perfectly usable for local coding through OpenCode. Verdict: very nice local and free backup to my usual GPT/Claude/DeepSeek for code. Good for Web searches via Brave search tool calls. What more do you want from a small local model?
- terhechte 2mo agoWhich setup did you use? MLX/GGUF, Quant, Engine (e.g. llama.cpp or MTPLX, etc)? There’s so much variety these days.
- dexterlagan 2mo agoIt was in LMStudio (llama.cpp), Q4 by Unsloth. Applied the recommended defaults published by Unsloth.
- cjbprime 2mo ago> General knowledge: I usually ask 2 questions many small models get wrong: summarize Operation Trojan Horse by John Keel, give publication year. Summarize the Ariel school incident of 1994. Qwen3.8 got the first question right, along correct publication year, gave glorious detail on Keel's theory, but got the second wrong. It thought that school was located in the USA. Ah well. I'm not sure that this means anything. You're asking a ~27GB file to have losslessly compressed the entire training set (which apparently is a large chunk of the entire internet). That's not possible. Whether it happened to encode these particularly obscure facts losslessly or vaguely isn't really telling you anything about how good a model it is.
- SamInTheShell 2mo agoIt's worth running, even quantized. I liked Meta Muse Glimmer's outputs, but qwen3.8-27b@q4_k_s kinda seems way better. Haiku/Sonnet kinda pairing in workflows?
- jamesblonde 2mo agoFor the full weights, unoptimized on vLLM with 2 Nvidia 6000 RTX 48GBs connected by NVLink, i only get 14 tokens/sec with open-code. For batched operations, it climbs to 55 tokens/sec. For FP8, on a single Nvidia 6000 RTX 48GB, i get 13 tokens/sec on a single GPU and 46 tokens/sec batched.
- nullc 1mo agoOn 2x RTX A6000 non-nvlink connected but communicating across the CPU, with llama cpp I get ~60 tok/s for Qwen3.8-27B-UD-Q8_K_XL without any batching.
- jishnuck26 2mo agowww.asuralist.in Couldn’t afford claude pro so I built an web based DSA coach that coaches you on DSA and System design in a socratic way. It uses a qwen 1.5B coder model and inferencing is all done on a CPU. ( who needs a GPU anyway )
- zyvop1 2mo ago[dead]
- sams99 2mo ago[dead]
- bricss 1mo agoBarely ~3 t/s on a laptop GPU > . <
- deleted 1mo ago[deleted]
- c16 1mo agoBig thank you to the Qwen team. 3.6 A3B was shocking good, and now I'm hoping they release an 3.8 A3B model too. Edit: Having used qwen3.8:27b-mlx on MBP M4 64GB, I get around ~45 tok/s. A3B would be great for smaller devices, but it's definitely usable. As I understand it it's a mixture of MLX and MTP.
- cogman10 1mo agoLooks like their A3B model is on the way [1] https://www.reddit.com/r/LocalLLaMA/comments/1voxppd/qwen_38_35ba3b_spotted/ https://www.reddit.com/r/LocalLLaMA/comments/1voxppd/qwen_38...
- brcmthrowaway 1mo agoThats a huge tok/sec. Prompt prefill is the bottleneck
- dofm 1mo agoSo far what I am seeing in my seemingly simple "Wordpress last-login plugin" test is that in xhigh reasoning mode (the default, seemingly) it overthinks so badly that it writes terrible bushy code with edge cases caused by going down rabbit holes. In "medium" reasoning mode, you get the classic Qwen wait/actually thinking loops you see in 3.6 that I guess will need to be interrupted in the way others do already with an over-thinking guard proxy. (In one of my test runs it is now on "OK TRULY FINAL APPROACH" after having got through "FINAL FINAL APPROACH". Can relate) It gets stuck in a thinking loop regarding the WordPress API and (resolvable) ambiguity in my prompt, that I guess might be resolvable with a custom skill with hints on how to look it up (and maybe with the devdocs MCP). In Low reasoning effort mode it flies through the task and writes pretty solid code. So maybe it is me overthinking what is needed here...
- nubg 1mo agohow do gemma4 or muse perform?
- dofm 1mo agoGemma 4 26B does really well at this specific task (and a general MySQL-related puzzle I test on). I rather like it and now they have fixed tool calling, I would use it. I think maybe it has been trained well with "consumer" programming languages like PHP that are sort of commonplace things people want to do. I think for less commonplace programming languages, maybe it's worse. Muse Glimmer thinks well and codes well in my tests; it does fine at this. I really like it so far, but my tests are fairly shallow. One thing I have been struck by — my prompt includes this sentence: "Please read the following and then ask me any further clarifying questions you need before proceeding with code generation." Almost all models I've tested interpret this as an instruction to ask questions regardless. Qwen 3.8 27B is the only one that either expresses confidence that it doesn't need to ask clarifying questions, or in higher reasoning effort ultimately asks questions, but offers up defaults I can choose with a simple reply.
- RobertasTa 1mo ago[flagged]
- XCSme 1mo agoIt's an amazing model, it's GLM-5.2 level[0], running locally... I tested it on my 3090, took like 8 hours to benchmark it and my room became a furnace (35+ deg outside temp), but it's really good. Now, in theory, you can talk directly to your computer and tell it what to do, and it does everything locally. [0]: https://aibenchy.com/compare/qwen-qwen3-8-27b-medium/z-ai-glm-5-2-high/ https://aibenchy.com/compare/qwen-qwen3-8-27b-medium/z-ai-gl...
- scirob 1mo agoWe maintain German Langauge index as no one publishes or reruns these sepeartly. Qwen 3.8 27B is a small improvement with some regressions in our benchmarks not a huge jump like benchmarks listed. https://dach.peerbench.ai/compare?models=qwen%2Fqwen3.8-27b,qwen%2Fqwen3.6-27b https://dach.peerbench.ai/compare?models=qwen%2Fqwen3.8-27b,... German language has never been a big focus for asian models but they still outperform Gemma models https://dach.peerbench.ai/compare?models=openai%2FQwen%2FQwen3.8-27B,openai%2FQwen%2FQwen3.6-27B,google%2Fgemma-4-31b-it https://dach.peerbench.ai/compare?models=openai%2FQwen%2FQwe... So in production we have been using Gemini Flash Lite as primary and fall back to Qwen when gemini servers are overloaded or just giving us 429
- _ache_ 1mo agoFrom your benchmark, Qwen3.8 is nearer than Opus 4.8 than Qwen3.6. 0.1pp but still. Also, a lot of people don't really care about german language capacity, maybe people programming in DDP idk. PS: You benchmark seems saturated. Most values sit @>75% in a benchmark generally indicate that it's no longer as useful as a <70% one. I mean, Qwen3.8 is 77.5% and Fable5 80%, the poll of values is from 65% to 90%.
- mixermachine 1mo agoDid you already test TranslateGemma? I use this model for my Android Studio Translation Plugin (https://plugins.jetbrains.com/plugin/30265-localizepipe https://plugins.jetbrains.com/plugin/30265-localizepipe) and so far it produces great results for its size. If there are other models (of similar size) out there, that are better at this, please let me know.
- cakbeslik 1mo agosmaller and smarter is always better
- jtrn 1mo agoI have been running it on my M5 Mac and was impressed with how well it worked with Pi coder. It can genuinely work as an assistant fully locally. It helps me configure Dockerfiles, fixed a couple of errors in a test Nuxt app, and so forth. Not very fast at 20 tps (8-bit quant for total memory usage around 30 GB), but enough to feel that I have a true local coding buddy. Then came the cold water shower. The agent kept trying to figure out a Nuxt icon package issue and was working on it. On the positive side, it was making steady and slow progress without getting stuck in doom loops. But after 20 minutes, I decided to test with Luna. So I switched in Pi and asked it to review the problem and fix it. Same session. Thirty seconds later, it was fully fixed. API cost on open router was $0.02, probably most of it due to the inheritance of the previous session. At that rate, the power consumption for local would be FAR higher than the API cost to solve the task. I wish it wasn’t so, but the cost per intelligence is just off the charts now with Luna. Now I am really liking that GLM 5.3 will probably run fine on 4x DGX Spark. If nothing else, the local models are truly usable for basic coding and assistance. I would have been blown away by the support I could have gotten with Qwen 3.8 when I was starting out coding. Hopefully, the local models will catch up AND the hardware becomes affordable in the future. Local models are keeping the largest LLM providers on their toes. But right now, it does not make economic or capability sense to run locally. It does make privacy, security, and vendor lock prevention sense, though.
- zerologadmin 1mo ago[flagged]
- kaycrafter 1mo ago[flagged]
- mmeyerlein 1mo ago[dead]
- deleted 1mo ago[deleted]
- prabhanjana_c 1mo agoOn my RTX3060 - 12GB VRAM + 24GB RAM , with below command ollama run qwen3.8:27b --verbose "explain mmap”, I got 2.41 Tokens/s, Not sure if that can be improved considering VRAM doesn’t fit the entire, model. Additional details: total duration: 8m18.2870918s load duration: 612.105ms prompt eval count: 12 token(s) prompt eval duration: 2.900965s prompt eval rate: 4.14 tokens/s eval count: 1193 token(s) eval duration: 8m14.660618s eval rate: 2.41 tokens/s System spec: NVIDIA GeForce RTX3060 AMD Ryzen 5 1600 Six-Core B450 AORUS M Mother board. NVIDIA-SMI 620.02 Driver:620.02, CUDA Version: 13.2
- anotherCodder 1mo agovllm on 4x 5090 is getting ~20 tok/s with mtp on (their own thread on the hf card). i had qwen3.8-27b up the day after release, one rtx pro 6000, 140 tok/s spec, 0.156s first token, full 262k. image and video on the same api. numbers: https://github.com/avifenesh/memra https://github.com/avifenesh/memra try it: https://inference.tiyuvta.ai/app https://inference.tiyuvta.ai/app $0.38 in / $0.20 cache / $2.60 out. openrouter's only host right now is 23 tok/s at $0.45 / $3.20.
- anotherCodder 1mo ago[dead]
- anotherCodder 1mo agoprice update: $0.40 in / $0.10 cache / $2.90 out