10 ms·
Qwen3.8-2.4T
https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B-FP8 https://huggingface.co/Qwen/Qwen3.8-2.4T-A95B-FP8
- volf_ 2mo agoA ~5TB model.
- PunchTornado 2mo agoThe card looks almost too good to be true
- guardiangod 2mo agohttps://unsloth.ai/docs/models/qwen3.8 https://unsloth.ai/docs/models/qwen3.8 The 1bit quant model is at an astonishing 397GB with 95B active per MOE. This literally puts Opus 4.5 performance level into a machine a normal person could buy, and still gets usable tokens/second. The full lossless model BF16 is clocking at 4.9TB. The model card claims the model to be between Opus 4.8 and Fable 5. Again that's astonishing as getting a machine with 7TB RAM (with context + KV cache) is still within the realm of medium size companies. Bad things: The open source version has its vision capability removed, and the context capped at 250k . I expect someone to bolt a Kimi 2.6 vision tower to it to restore the vision capability (at less performance of course). For context, I played around with extending the context to 600k for Qwen 3.5 397b, and the context remained stable up to around 480k. It'd be interesting to see if the same can be done to Q3.8 . Also no out of the box DSpark/DFlash support. MTP is present so we should at least get some boost in TP speed.
- ilc 2mo agoTo compare a 1 bit quant to the full fat model is misleading. Honestly this model people at home can tinker with, if you have a big enough Mac. Maybe 4 Strix Halo/DGX Spark, and then at 1 bit quant? Nah. Use the right sized model, for your hardware. You'll get better results.
- guardiangod 2mo agoExtremely large 1 bit models are usually within 50-60% of KV divergence to lossless models. In this case I think the comparison to Opus 4.5 is a fair assessment. Extremely large models don't suffer as much from quantization due to its weight topology also contains encoded information, so the loss of info from any one weight is somewhat mitigated.
- ilc 2mo agoAny one weight, but all of them. And also crushing the architecture itself? I wouldn't pick up 400gb of hardware to run in that mode. I might try it for fun, but even then you are looking at handling a 95GB active parameter set. This is NOT a model for most home labs. I'm sure some can and will use it. But most, should steer clear.
- kadoban 2mo agoI wouldn't just rush out and buy hardware, but there will be benchmarks after a while to make an informed decision. 95GB active is not _too_ bad, would require some creativity and $$, but I bet I could do that at home for less than a cheap car.
- dist-epoch 2mo agoKL divergence (you misspelled it) doesn't tell you anything about capability drop - how much did this particular benchmark (thus ranking among models) change after 10% or 50% KL divergence?
- frotaur 2mo agoNot only that, but KL divergence is not a '%'. It's just a number ranging from 0 to infinity that tells you the 'distance' between two probability distributions.
- auspiv 2mo agoOpus 4.5 level of performance is also accessible with deepseek-v4-flash-0731 (0731 being the july 31 update) which is much, much, much smaller. 2x RTX pro 6000 blackwell can run it. 4x can run it very comfortably
- Philpax 2mo agoWhat do you need the extra 2 for? Tensor parallelism?
- arjie 2mo agoLonger context and more cache. The problem is that native format with DSpark enabled you have very little room on the VRAM.
- Philpax 2mo agoI was under the impression that you could fit the full 1M context within the 192GB VRAM as a result of DeepSeek's various architectural advancements, but I'll grant that DSpark + a larger pool for concurrency may necessitate more VRAM, yes.
- guardiangod 2mo agoI am running DS v4 flash 0731 lossless at 80t/s right now. It really is not at Opus 4.5 level (for my workload). I would say it's around 3.7 Sonnet, which is still pretty good, but other models such as GLM 5.2 are still leaps better. Of course I run DSv4 flash over GLM 5.2 for a few very good reasons, but intelligence is not 1 of them.
- MrDrMcCoy 2mo agoDespite fitting into VRAM, I can't get DSV4 to run at usable speeds on my AMD hardware. The upcoming qwen3.8 27b greatly excites me, and I hope it can outperform Stepfun 3.7 Flash, which is the best thing I can run today.
- kennywinker 2mo ago
- pil0u 2mo agoI don't understand the logic behind model sizes and quantization. Suppose I have 100GB of unified memory, how should I know which model suits it best? I understand how a 2.4T model wouldn't fit, but I don't understand the impact of quantization and whether I should use a 200G model quantised to fit say 90GB of memory, or a non-quantised 90G model.
- markasoftware 2mo agoThere's no rhyme or reason to it. Quants aren't benchmarked much. Generally 4bit better than smaller model 8bit
- NitpickLawyer 2mo agoIt really depends. It used to be easier to have a rule of thumb, but now it's not clear anymore. Now there are a lot of things to consider, such as a model's kv efficiency (how much context you can fit), MoE v. dense, QAT or not (Quant aware training) and so on. The old rule of thumb was that a lower quant of a larger model > higher quant of a smaller model. That being said, for some things going lower than fp8 will see a lot of degradation in generation quality. Except if the model comes with QAT 4bit quants. Then there's also nvfp4 w/ calibration data, which also can improve things. So it's really not easy to tell "at a glance" you'd have to test them yourself on your hardware.
- onlyrealcuzzo 2mo agoStandard models are designed to quantize down to 4-bits relatively well. Anything below that, and especially 1.58b - is typically complete garbage, and you're much better off running a model 100x smaller at regular precision (compared to one 7x smaller quantized into complete garbage). If the model was designed specifically to quantize down to 1.58b, then it's different. AFAIK, there's no large models designed for this yet.
- richardfey 2mo ago> If the model was designed specifically to quantize down to 1.58b, then it's different. > AFAIK, there's no large models designed for this yet. Isn't BitNet b1.58 2B4T what you are looking for? (haven't tried it myself though)
- ekianjo 2mo ago> The 1bit quant model i at this kind of quantization is it useful though?
- wolttam 2mo agoOpus 4.5, even 4.6-level performance has been around since July 31st in 284B total params and just 160GB of weights at native FP4 quantization- DSv4 Flash.
- simonw 2mo agoIs this the largest ever open weight model release by parameter count? I think it is.
- NitpickLawyer 2mo agoNo, Kimi k3 is 2.8T params. This is 2.4T params but ~5TB weights because it was released in bf16 and ~2.5TB for the fp8 version. Kimi k3 launched with QAT 4bit, so ~1.5TB weights.
- Mercuriusdream 2mo agoKIMI K3 was the biggest open weight release afaik; It is 2.8T-A100B if I'm correct
- NitpickLawyer 2mo agoSupposedly this is a Kimi k3 rival. Bit of a chonker, especially since they only released bf16 and fp8. So at launch this will be harder to serve than k3. No QAT on q4 means that someone with deep pockets (nvda?) will have to quant it, with plenty of calibration data. Should bring it ~1.3TB, so around k3 size. License pretty similar to k3 with some caveats. Free to use for internal or <50M$ revenue / year. Limitations above that threshold for serving the model or services targeting coding / productivity agents. Benchmarks are looking good, trading blows w/ opus4.8 and sol, generally 10-20p under fable. But that's neither here nor there w/ qwen, their benchmark to real world usage correlation has been iffy in the past. The local model 3.8-27B announced for Friday, same time so ~48 hours from now. That'll be a bit more exciting for a lot more people, since 3.6 was quite good for local inference, and their 3.7-max -> 3.8-max shows a lot of improvement.
- MrDrMcCoy 2mo agoLlama.cpp can quantize without special training, but I'm not sure if any special model architecture support is needed to read it in the first place. If it can be converted to gguf at all and you know what tensors to target, it can get the full ternary bonsai treatment today.
- binary132 2mo agoQAT is an optimizing quantization algorithm, not naive quant.
- MrDrMcCoy 2mo agoRight, but the way they phrased it suggested that without QAT it could not be quanted at all.
- dragonwriter 2mo agoIsn’t QAT a training approach (roughly, simulating quantization in the forward pass during training so that quantization of the level targeted in training has close-to-optimal behavior), not a quantization algorithm? Hence, the name?
- deleted 2mo ago[deleted]
- l72 2mo ago> In particular, Qwen3.8-Max is the official version based on Qwen3.8-2.4T-A95B with more features, such as vision input & non-thinking support, 1M context length by default, official built-in tools, etc. That is unfortunate, that the open weight model doesn't have vision support or the 1M context length...
- wren6991 2mo agoPeople have had surprising success adding vision to open-weight LLMs that ship without it, like DSV4 Flash [1] or GLM-5.2 [2]. Given this model is already vision-trained I expect that approach will work well here. [1] https://old.reddit.com/r/LocalLLaMA/comments/1vl6ior/i_gave_deepseek_v4_flash_basic_vision_by_training/ https://old.reddit.com/r/LocalLLaMA/comments/1vl6ior/i_gave_... [2] https://huggingface.co/baseten/GLM-5.2-Vision-NVFP4 https://huggingface.co/baseten/GLM-5.2-Vision-NVFP4
- mips_avatar 2mo agoQwen3.5 was awesome: fairly open and fully featured. 3.8 lacking vision, nerfing thinking modes, and low context length feels pointless.
- theanonymousone 2mo agoDo we know if AA and DeepSWE benchmarks are on bf16 or fp8 quantisations?
- dhx 2mo agoAlso of interest: DeepSeek V4-Pro-0813 (1.6T-A49B) benchmark scores have apparently just been announced on the DeepSeek WeChat channel and they're sitting about Fable 5 level.[1] [1] https://www.reddit.com/r/LocalLLaMA/comments/1vmi0fg/deepseek_v4pro0813_benchmarks/ https://www.reddit.com/r/LocalLLaMA/comments/1vmi0fg/deepsee...
- svantana 2mo agoAnd relatedly, just now available on OpenRouter https://openrouter.ai/deepseek/deepseek-v4-pro-0813 https://openrouter.ai/deepseek/deepseek-v4-pro-0813
- onlyrealcuzzo 2mo agoIsn't this quite a bit behind Sol and Fable and even ChatGPT 5.5 xhigh and Opus 5 max? In terms of what you get for what you pay for, it's incredible - probably by far the best. But unless I'm reading things wrong, it does not appear to be top-of-the-line.
- neosat 2mo agoThis may not 'quite a bit behind' those at all. If you look at the benchmark numbers they are very comparable to Fable, but beyond a certain point the benchmark numbers don't tell you much. Opus #5 beats Fable on some benchmarks but given similar cost almost everyone who has used those two models will prefer to use Fable. At this price range $0.87per 1M they will get a lot of usage of people trying it out. Given the benchmark numbers, for many people and many use cases this will become their primary driver. There are people and use cases where Fable, Sol will work better but those are likely not the target of DeepSeek anyway. In terms of performance and price pareto curve I don't think any model can beat this today (though openAI is doing some exciting recent work in efficiency) - which is a remarkable feat for the DeepSeek team. Either way, what a time for consumers of these models :)
- SwellJoe 2mo agoI stopped picking Fable because it refuses based on guardrails so often. I do a lot of security related work, and Fable just won't do any of it. So, I don't bother. Unfortunately, Opus 5 also refuses quite a bit of security work, now, as well, so my Anthropic subscription becomes less useful by the day. Fable may be better, but if it won't do the work... DeepSeek and Kimi K3 will happily do security work, and they do it pretty well.
- jephs 2mo ago"QwenSVGBench" elo 1713, pelicanmaxxxing confirmed?
- ByteWarden 2mo agoMore curious about how qwen3.8-27B performs. That's the size that I can run locally.
- gilgoomesh 2mo agoYeah, I must have misread the press release last week as I thought it would be released at the same time.
- drey08 2mo agoReleases tomorrow: https://huggingface.co/Qwen/Qwen3.8-27B https://huggingface.co/Qwen/Qwen3.8-27B
- ycui7 2mo agothe a little disappointing part is this is released in BF16. so i suppose no QAT was implemented.
- cautiouscat 2mo agoI've been wanting to run open weight models lately to give them a shot with OpenCode. However, I get the impression that models like Qwen and Kimi k3 are impossible to run locally? I have a RTX 5090 and 64 GB of RAM but the models seem to be much larger than that. What's the route to start using these models? Bedrock?
- zeeveener 2mo agoYou could easily run any of their 30B-or-less models which is what most people are waiting for. Apparently the ~30B variant will be released on Friday?
- daemonologist 2mo agoBedrock seems to have stopped adding new open-weights models, and mostly only has Anthropic and OpenAI stuff now. You can get Qwen 3.8 directly from Alibaba: https://www.qwencloud.com https://www.qwencloud.com (proprietary variant) or from DigitalOcean (this variant, probably also from others soon). On your 5090 you could easily run a smaller model like Qwen 3.6 27B: https://huggingface.co/collections/Qwen/qwen36 https://huggingface.co/collections/Qwen/qwen36 or Gemma 4 etc., or as mentioned there's a Qwen 3.8 27B coming out in a few days.
- cautiouscat 2mo agoWhat does the number before the B signify?
- CamperBob2 2mo agoNot seeing the upside versus K3 here, especially with the intentional capability loss. Read the room, Qwen. It's not a good time to hobble your releases.
- deleted 2mo ago[deleted]
- imagetic 2mo ago3.8-27B LETS GO! best crypto-bro impression I can do...
- frozenseven 2mo agoVocabulary size ~248k. A bit bigger than other recent Chinese models (Kimi K3 ~164k, DeepSeek-V4 ~129k, and GLM-5.2 ~155k). Make of this what you will.
- miohtama 2mo agoDoes this mean its tokenizer is somehow tuned?
- Ey7NFZ3P0nzAe 2mo ago> Make of this what you will. I'm interested in your take on it. IIRC Gemma family models too have a ~250k vocabulary size
- octocop 2mo agowhen will we see MIT license Qwen again?
- boutell 2mo agoI'll just fire that up on my Intel n100...
- XCSme 2mo agoThat's a really cool hamster [0], unfortunately it's really expensive now, 2x more expensive than Grok 4.6[1]. [0]: https://aibenchy.com/compare/x-ai-grok-4-6-high/bytedance-seed-seed-2-1-turbo-low/qwen-qwen3-8-2-4t-a95b-low/bytedance-seed-seed-2-0-code-low/#showcase=9f4b661ee684c1c2 https://aibenchy.com/compare/x-ai-grok-4-6-high/bytedance-se... [1]: https://aibenchy.com/compare/x-ai-grok-4-6-high/bytedance-seed-seed-2-1-turbo-low/qwen-qwen3-8-2-4t-a95b-low/bytedance-seed-seed-2-0-code-low/ https://aibenchy.com/compare/x-ai-grok-4-6-high/bytedance-se...
- XCSme 2mo agoInterestingly, the high variant does a lot worse and failed to generate a valid SVG (and the low variant use more tokens than the high one, so maybe their reasoning efforts are not working properly). The solar system animation is also the coolest looking I've seen, unfortunately the animation doesn't work: https://aibenchy.com/compare/qwen-qwen3-8-2-4t-a95b-low/qwen-qwen3-8-2-4t-a95b-high/x-ai-grok-4-6-high/?showcase=solar-system-css#showcase=c098d6ee18b748dd https://aibenchy.com/compare/qwen-qwen3-8-2-4t-a95b-low/qwen...
- yakisuragi 2mo agowhy is this page suddenly 404? Is Alibaba going back on their words? https://modelscope.cn/models/Qwen/Qwen3.8-27B https://modelscope.cn/models/Qwen/Qwen3.8-27B
- Bob_bo 2mo agoPeople online say this model's performance isn't very good; what do you think of it after using it?
- edg5000 2mo agoShall we bet on when the hardware needed for this (without quantizing and at good speed) will reach < 10k USD? I'm betting 2040. I can download it now, and then get the hardware later. Eventually we can all have these things running 24/7 in our home if we wanted to. I currently would not have any task for it that would really utilize the hardware 24/7, but maybe in 20 years I will.
- arthurcolle 2mo agoapproximately $20 million for 750TB unified memory custom interconnect right now
- edg5000 2mo agoWow, that's way higher than I assumed. I looked into it and remember ending up with something like 200k, but I must have been off. That's crazy.
- arthurcolle 2mo ago[dead]
- ak_t 2mo agoI think it is more likely that a smaller model (<400B) with similar intelligence gets developed long before the hardware to serve a 2.4T model gets cheaper than 10k.
- edg5000 2mo agoAh, I hadn't thought of that. To what degree have we seen this already? What would you consider as the biggest jump in intelligence per weight?
- ak_t 2mo agoIt's been pretty consistent, the smaller models (hundreds of billions of params) usually catch up in 6 months or so to their frontier counterparts, at least on benchmarks.
- jannishan 2mo agoHas anyone compared the programming capabilities of Qwe3.8 and Kimi K3? Which one is better?
- bunkydoo 2mo ago[dead]