8 ms·
Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)
- tarruda 1mo agoCan you share the source for the parameter count (125B A6B)? I didn't see it anywhere in the page.
- petu 1mo agoIt was in description under the countdown initially, but was quickly removed. It also said 51B of n-grams and new attention (IIRC it said "Qwen Sparse Attention"). edit: here's a random screenshot https://x.com/AiBattle_/status/2092210011858460819/photo/1 https://x.com/AiBattle_/status/2092210011858460819/photo/1
- NitpickLawyer 1mo agoThis is what I copied from the en version of the modelscope page, right when they published it: > Redisgned Multimodal MoE Model: 125B main model parameters, supplemented by an additional 51B N-gram embeddings,and 6B parameters activated per token. > Efficient Training and Inference: Significantly reduces training and inference costs. At ~1/9th the training cost,Qwen3.8-Flash-Next achieves comparable capability against Qwen3.7-Plus, while being more capable in areas of coding and cowork. There was another paragraph about a new attention, but I didn't copy that.
- fcanesin 1mo agoHF link: https://huggingface.co/Qwen/Qwen3.8-Flash-Next https://huggingface.co/Qwen/Qwen3.8-Flash-Next
- honestlyranked 1mo agoAlibaba is giving sleepless nights to the tech giants
- WithinReason 1mo agoSounds like a line from a fairy tale
- drannex 1mo agoTo be fair, Alibaba IS a tech giant, one of the biggest in fact. They are just giving sleepless nights to the western tech giants.
- _ache_ 1mo agoI'm hearing Tencent, Zhipu and Baidu shaking from here. It's fair to assume BATX / 6 Tigers don't sleep very well either.
- mrdoe 1mo agolol blocked with dns4eu what a joke this resolver has become
- cogman10 1mo agoWow. I wasn't expecting this. I thought they were going to do a 35B model instead.
- hasteg 1mo agoAs a 5090 owner and local model enthusiast, I was hoping it would be 35B A3B so I could run it myself =(.
- cpburns2009 1mo agoYou can run the 27B released last week. I haven't tried it yet myself but the 3.6 version runs great on my 5090.
- Philpax 1mo agoStrongly recommend https://github.com/Neroued/ninfer https://github.com/Neroued/ninfer, which can pull ~180 TPS on 5090 with 3.8, and 500 (!) with 3.6 35B-A3B.
- cpburns2009 1mo agoI've been waiting for the dust to settle on this model so I can find a good runtime setup. I'm definitely bookmarking this. Thanks!
- hasteg 1mo agoI've been running 27B a lot, I am honestly shocked at how well it performs. It's mind blowing how well the small qwen models (and particular, the 3.8 model) runs locally. It can create some really impressive toy coding projects. I haven't really done too much integrating into my actual workflows because I pay $100 for Claude Max, but I can see it being pretty helpful in that.
- Tuna-Fish 1mo agoThe 27B one is great on a 5090. This one is basically aimed at macs, Strix halo and DGX Spark.
- big-chungus4 1mo agoI hope there is going to be a free endpoint... Unlike 35B-A3B, I am nowhere close to running it locally
- blurbleblurble 1mo agogg
- david927 1mo agoWell put and succinctly put. And if OxA is a flash model? it becomes: goodnight
- big-chungus4 1mo ago> We are releasing these architectural improvements ahead of time so that the community can prepare for the upcoming full family of Qwen4 models. That gives me hope that "full family" means it will include smaller models like 4B.
- culi 1mo agoWhat are the use-cases for a model as small as 4B?
- big-chungus4 1mo agoThey are great base models for fine tuning on both text and visual tasks. Many OCR and object grounding models are based on small Qwen models, though they often replace vision encoder with a bigger one. Qwen3 5-4B is the biggest model I can find tune in my laptop. And when I upgraded the model from Qwen3-4B to Qwen3.5-4B, both vanilla and fine tuned performances jumped significantly on a classification task. Those models are great when you have very little data or very low diversity of examples, where it's not possible to train a neural net from scratch as it will just memorize the data. The best you can do is fine tune a generalist model that can already do the task for small number of steps until it starts over-fitting, or on some cases you can do even better though RL.
- pwython 1mo agoI was already rolling around the idea of a 128GB M5 Max MBP. Now this! A 4-bit MLX quant with 128k window should fit perfectly, in the 50-70 tok/s range.
- sscaryterry 1mo agoI have a 128GB M5 Max, and it sucks at this stage. 50-70 tok/s might be something...
- smcleod 1mo ago50-70tk/s is what I get on my m5 max on a 5-6bit Qwen 3.8 27B?
- Casteil 1mo agoI don't know what black magic you're up to but I see more like 30-35t/s on a 16" M5 Max using 3.8:27b Q4, regardless of whether it's mlx or gguf. qwen3.5:122b-a10b is significantly faster at around 60-65.
- syntaxing 1mo agoWith MTP? I get 25-30 TPS on a strix halo. 50+ on a M5 max should very doable. Dflash (2) will push your TG even further
- Casteil 1mo agoIt's a bit deceptive to state inference speeds without mentioning the additional things you're doing to achieve them
- smcleod 1mo agoNo magic, just oMLX with MTP. You can look through the speed the community is getting here: https://omlx.ai/benchmarks/performance?model=qwen3.8&chip=&chip_full=M5%7CMax%7C40&quantization=&context=&pp_min=&tg_min=&sort=tg_tps&order=desc https://omlx.ai/benchmarks/performance?model=qwen3.8&chip=&c...
- tw1984 1mo agoQwen4 sounds exciting
- luciandan 1mo agoComing soon...
- ddtaylor 1mo agoI enjoy the Qwen models a lot, but building things on top of them with OpenRouter has been painful. OpenRouter does a lot of great work and I really enjoy being able to use different models so easily. I like when a provider is phasing out an older model that still works for my needs and the price is much lower. It seems like such a good win-win. However, the problem is that many Qwen models have almost no capacity or is so flaky you literally have to just litter your code with a blacklist/whitelist of providers. OpenRouter has some attempts to solve this, but they don't work. In fact, OpenRouter has a lot of really cool stuff that is documented, but if you read the code it's not yet implemented or isn't actually there yet, which is a shame. I tried to get in contact with them at OpenRouter about this and I was interested in working with them in the past, but it's difficult to get in touch with the right people and they are growing very fast. I expect being acquired by Stripe will accelerate those problems in some ways. I have no doubt they will resolve all of these issues eventually and scaling that much that quickly is really hard, so kudos to them, but the road has been pretty lame and taken some wind out of my sails.
- irthomasthomas 1mo agoOpenrouter was pretty great before prompt caching became common. Now it is extremely expensive for most individual workflows, unless you spend a lot of work customizing router preferences, and then you still get a worse cache hit rate than using the provider directly. I only keep $5-$10 in OR for occasional testing.
- dackdel 1mo agodidnt stripe acquire open router? so i assumed its sunset.
- copperx 1mo agoWhy is caching affected when not using the provider directly?
- irthomasthomas 1mo agoI don't know. But take a look at https://openrouter.ai/deepseek/deepseek-v4-pro-0813#pricing https://openrouter.ai/deepseek/deepseek-v4-pro-0813#pricing for instance, where the deepseek provider shows an 85% cache hit rate, while the same one on zenmux is 98%.
- BrucecarlL 1mo agoWaiting for the performance report! Ai hope it can beat DS
- bellowsgulch 1mo agoReally happy for those with 128GB+ RAM. Sitting here with my Apple M1 Max with 64GB though. Was looking forward to a Qwen3.8-35B-A3B like many others.
- dofm 1mo agoHave you tested Muse Glimmer in low reasoning strength? Token generation is slow (and prefill is) but you will likely find it solves actual problems faster than Qwen 3.6 35B-A3B.
- bellowsgulch 1mo agoI’ll give it a try! Thanks for the heads up!
- dofm 1mo agoI’m using the Unsloth 4-bit quant. To change the reasoning strength you just put text in the system prompt. From memory it is: Reasoning strength: low
- metrofun 1mo ago[dead]
- hedora 1mo agoTime to dust off my 128GB strix halo (literally—it’s been dusty, and it’s running a bit warm these days). Any idea where this model sits according toquality benchmarks? Pre-bubble MSRP on this hardware was $1400, and it draws 200-ish watts, putting it down into consumer territory. I’m wondering if it can replace claude for llm-friendly coding tasks.
- cpburns2009 1mo agoSo back in the Qwen 3.5 release, the 122B-A10B model scored slightly better than the 27B model. I'd expect this new 125B-A6B to perform similarly to the recently released 27B. Qwen3.8 27B is supposed to rival Sonnet/Opus 4.6.
- hedora 1mo agoThanks. My current stack ranking of anthropic models is: 4.6 ~= 4.8 4.7 much worse. Fable and newer consistently tells me to pound sand, so I’m not sure what I’m paying $200/month for. 4.8 sometimes does too, but it’s at least usable most of the time. So, I’d expect this to mostly replace Claude for my workflows. The main tradeoff for me should mostly be token throughput vs. no longer really trusting anthropic.
- cyanydeez 1mo agoI've got the A10B hooked up to deer-flow and it does remarkable well when you dont need to baby sit it.
- hugmynutus 1mo agoQwen3.8/Qwen3.6 has a weird self doubt/thinking too much problem. You can prompt it away. I would say it "approximates" Opus 4.X class models well enough especially for coding/linux problems. The only reason I stopped using it as much is I was getting 25-35tok/s on Intel B70 (non-quant) which made some responses slow. For a long running/autonomous task, it would probably be sufficient.
- eightysixfour 1mo ago
- notnullorvoid 1mo agoIt will be interesting to see the intersection of this with inference engines like FreeToken which improve distribution of work for MoE models across CPU/RAM and GPU/VRAM. If all it takes for a competitive model to run locally at good speeds is a used 3090 and some DDR4, then we might be in for the year of local AI. https://github.com/FlashML-org/FreeToken https://github.com/FlashML-org/FreeToken
- Zylokloto 1mo agoYou can already run it locally its just not the same. It is still slow, a lot slower than what you are used to with claude and co. And as soon as you increase context size, your memory requirements jump. Then when it runs for 30 minutes for something claude needs 5, your device will get hot. And even a used 3090 is apparently now between 1-2k.
- notnullorvoid 1mo ago> It is still slow, a lot slower than what you are used to with claude and co. That really depends on the model, I run a few models locally. All at speeds comparable to or faster than Opus. In general we haven't reached the ceiling for what performance we can get out of consumer hardware. As evidence by FreeToken which hasn't even added MTP/speculative drafting support yet, which will add another boost. > Then when it runs for 30 minutes for something claude needs 5, your device will get hot. I doubt the timing differential here, but even still I run my 3090 pretty heavily with inference workloads and it stays cooler than when I use it for gaming. > And even a used 3090 is apparently now between 1-2k. Yeah I guess the price went up significantly in the last couple months, used to be hovering around 1k. 3090 isn't the only option though.
- nitin7 1mo agoWhich models when run locally come close to Sol and Opus, from your experience? And which harness do you use?
- 1mo ago
- isatty 1mo agoCan I run a fp8 quant with 96gb VRAM?
- cpburns2009 1mo agoOnly VRAM? Unlikely unless you can also load the whole model into regular RAM. The previous 3.5 release was 250gb at BF16, so FP8 would likely be around 125gb. Your best best is FP4/Q4.
- isatty 1mo agoVery sad. I try not to go below q8.
- SwellJoe 1mo agoFinally, a reason to own a 128GB Strix Halo or GB10 device. Or a reason to consider the new Mac Studio. I have a Strix Halo and dual 32GB GPUs in my desktop, and the latter is pretty much always better for running local models because it's quite a bit faster due to higher memory bandwidth. There simply haven't been any models that are better than Qwen 27B or Gemma 31B, which run comfortably in 64GB with big context. And, MoE should make it run at a close to usable speed.
- jubilanti 1mo agoA 3060ti 8gb, released in 2020, has 448 GB/s of bandwidth compared to the Halo 256 GB/s The 3080ti is 912.4 GB/s
- embedding-shape 1mo agoAnd the newly announced/launched Apple M6 has 170GB/s of unified memory bandwidth, meanwhile M5 Ultra gets 1.2TB/s of unified memory bandwidth. https://www.apple.com/newsroom/2026/08/apple-introduces-m6-and-m5-ultra-for-a-big-leap-in-performance-and-ai-compute/ https://www.apple.com/newsroom/2026/08/apple-introduces-m6-a... Not sure if the first one is a typo on their press release, can't be just 170GB/s then be pushed for AI use, can it? Could be a different measurement I suppose...
- jtbayly 1mo agoYou got me curious so I looked up the previous chips[0]. Memory bandwidth M1: 68 GB/s M2: 100 GB/s (47% increase) M3: 100 GB/s (0% increase) M4: 120 GB/s (20% increase) M5: 153 GB/s (27.5% increase) So, M6: 170 GB/s (11% increase) doesn’t seem impossible, though I would have expected more. [0]: https://www.jdhodges.com/blog/apple-cpu-compared-m1-m3-m3-m4-m5-max/ https://www.jdhodges.com/blog/apple-cpu-compared-m1-m3-m3-m4...
- mhast 1mo agoThe different models of chips and memory config have very different memory speeds as well. Eg the M4 Max 128GB has a bandwidth speed of 500GB/s+. And that's true for other models as well. But as you note, the base speed has also increased over the versions.
- onesandofgrain 1mo agowhere are the humans geez
- syntaxing 1mo agoReally looking forward to this, 27B is a struggle with a strix halo and Laguna 2.1 can do stupid things for tooling calls.
- cpburns2009 1mo agoYeah 27B is way too slow for the Strix Halo. Laguna was better but still slow when I tried it. Qwen3.6 35B is still the best today.
- deleted 1mo ago[deleted]
- cyanydeez 1mo agoit'll hopefully improve with more MoE and half the prefill/generation. I think it's the sweet spot for the strix halo for smarter or vibe tasks.
- puzzlingcaptcha 1mo agoWhat sort of pp/tg speed do you get on a Strix Halo?
- cpburns2009 1mo agoThis is the best I got, all with Unsloth's quantizations. Laguna-S-2.1:UD-Q4_K_XL (no MTP) pp=186.4 t/s tg=27.8 t/s Qwen3.6-35B:UD-Q4_K_XL (with MTP) pp=404.4 t/s tg=83.2 t/s Qwen3.6-27B:UD-Q4_K_XL (recorded pre-MTP) pp=343 t/s tg=12.1 t/s Laguna actually performed better than I remembered. I thought it was slower.
- deleted 1mo ago[deleted]
- htrp 1mo agoHave you benchmarked against full precision models for accuracy/ performance?
- dmead 1mo agoThis is great. I have a weird system layout (192gb system ram, 8gb vram). the mixture of experts models have been nice when i can run the dense reasoning layers on the gpu (which somehow fit?!) and then the expert on the cpu. its worked out to to 40 tokens/seconds on their 80b-a3b model. we'll see how much of a hit this is.
- cyanydeez 1mo agooooh, I like a6b; that will be nice. 3.5 A10B qwen works really well in deer-flow when you want to seriously vibe code or research and you're just not going to baby sit.
- Catloafdev 1mo agoVery curious to see how this compares to Deepseek v4 Flash. I have to assume they wouldn't be releasing this if it was worse.
- natrys 1mo agoWhy not? It's not really competing in the same size class. Besides, as they explicitly wrote here, the main goal for this release is not performance, rather to serve as a reference for inference runtimes about what to implement. So that later Qwen 4 can be released with zero day support.
- Catloafdev 1mo agoGood point, I didn't see that. I guess I categorized them in the same bucket of 'runs on 128gb machines' Guess Qwen 4 is the one to wait for.
- NitpickLawyer 1mo agoTheir "next" variants are usually undercooked, but useful for the community to verify support for inference stacks. This will likely be the same.
- c16 1mo ago+1 to the long list of people hoping for Qwen3.8-27b A3B.
- NitpickLawyer 1mo agoThey've said no moe for 3.8, and since they're already releasing a qwen4 early preview, they're probably focusing on that arch going forward.
- kamranjon 1mo agoWhere did they say that? My understanding of this 3.8-Flash-Next release is that it's a MOE (as per the title of the posting here, 125B a6b)
- NitpickLawyer 1mo agoA bit of context: 3.5 was the last version where they released their entire suite of models 2b-400b. Then 3.6 got a 27b dense and a 35b moe. Then 3.7 was API only, and 3.8 got only the 27b dense. The devs confirmed on twitter that 35b moe would not come. So that's what I meant by 3.8 is not getting a moe. 3.8 next is not really a 3.8 (but I guess they had to disambiguate from the previous next). It's a preview of qwen4 architecture (and it is an moe + ngram), released early as a preview, and to help the community sort out inference before qwen4 releases.
- WiSaGaN 1mo agoYou probably meant Qwen3.8-35B-A3B. But judging from some of the words from their team, it seems unlikely unfortunately.
- vorticalbox 1mo agoThey normally release a 35b dense and an 27b moe (4B active per token) For context 35B on my m4 runs at 10 tokens a second, 27B moe runs 50-60 tokens a second.
- lousken 1mo agogpt oss killer? this can easily run on a server cpu with its memory bandwidth
- fkndkfn 1mo agoI can feel Dario Amodei's tears in the announcement :)
- Alien1Being 1mo ago[flagged]
- latentsea 1mo agoSome of us are actually quite excited about this release given the performance of Qwen3.8-27B.
- Alien1Being 1mo ago[flagged]
- latentsea 1mo agoCare to expand on that?
- freddiehdxd 1mo agoDoes it support vision?
- nightfuryg 1mo agoyes
- system2 1mo agoThese companies are naming their products worse than I was naming my half-baked software in the 90s as a junior developer.
- vegnus 1mo agoI have an m1 max 64gb macbook. Anything I can do to get 3.8 27b at more than 10 tok/s or am I relegated? 3.6 a3b is good but its not as good
- _ache_ 1mo agoWhat will be the requirement, like 128G of RAM and 12G of VRAM ?
- embedding-shape 1mo agoHow long is a rope? Technically you could probably run it off a SSD, but it'll be slow as molasses. If you want it "fast", you want it all within GPU and VRAM, who knows what that'd be. If the engram parameters are separate, I guess it'd be like BF16 ~400 GB, FP8 ~200 GB, NVFP4 ~100GB. Otherwise maybe like ~300GB, ~150GB and ~70GB or alike, don't quote me that, only some guesses. The one who waits will see :)
- _ache_ 1mo agoI think a reasonable expectation of MAX requirement to claim "runable on consumer hardware" is to 32G VRAM and 128GB RAM and it run at +10tps.
- latentsea 1mo agoI've got an R9700 32GB and an RTX 5060 Ti 16GB plus 64GB of system ram. Hoping to be able to run this at around 30 t/s on a Q4 quant. Hoping. Really hoping. Anything below that isn't usable as a daily driver since at deep context it drops quite significantly, so if you start out at say 20 t/s then you'll wind up at like 10 t/s and 20 t/s is already too slow.
- embedding-shape 1mo agoOk, so you already know what your expectations of the requirements are, and you aren't interested in more conservative perspectives, why do you ask to begin with?
- _ache_ 1mo agoI will rephrase it. Will it run on any consumer hardware?
- deleted 1mo ago[deleted]
- vietvu 1mo agoThis is a gift!
- vikasgrac 1mo ago[dead]
- listingbott 1mo ago[dead]