3 ms·
I think everyone is hoping this! It would be great if they'd release an MoE model somewhere between the 35B size of 3.6 and the 122B version of 3.5 - it could
by mft_ 3mo ago
I think everyone is hoping this!
It would be great if they'd release an MoE model somewhere between the 35B size of 3.6 and the 122B version of 3.5 - it could be a great balance of speed and ability for people with reasonably powerful but not insane home computers.
- nsbk 3mo agoIndeed! That would be the sweet spot for my 2x3090 rig
- embedding-shape 3mo ago> the 122B version of 3.5 Yeah, this is what I'm holding out for, the NVFP4 variant of 3.5 122B is blazing fast with reasonable quality and even with max context fits perfectly within 96GB.
- nsbk 3mo agoSweet. Are you running a Mac Studio Ultra or 4x3090?
- embedding-shape 3mo agoSimpler (and faster): a single RTX Pro 6000 :)
- nsbk 3mo agoHahaha well played! Hat tip
- androiddrew 3mo agoCurious, do you find the 3.5 120B sized MoE works better than the dense 3.6 27B?
- embedding-shape 3mo agoYeah, Qwen3.5-122B-A10B-NVFP4 produces better responses than Qwen3.6-27B-NVFP4 (both from unsloth), but I'm mostly using them for programming in various ways, mostly Rust, Clojure, Python and JavaScript, and some translations tasks, but not much more than that, so YMMV. Edit: as a concrete example, I'm working on a "optimization framework via agent harness" right now, Qwen3.6-27B-NVFP4 is often unable to actually complete the optimization within 100 turns, while Qwen3.5-122B-A10B-NVFP4 has no issues finishing within ~50 turns or so.
- 3abiton 3mo ago> Yeah, Qwen3.5-122B-A10B-NVFP4 produces better responses than Qwen3.6-27B-NVFP4 (both from unsloth), but I'm mostly using them for programming in various ways, mostly Rust, Clojure, Python and JavaScript, and some translations tasks, but not much more than that, so YMMV. > > Edit: as a concrete example, I'm working on a "optimization framework via agent harness" right now, Qwen3.6-27B-NVFP4 is often unable to actually complete the optimization within 100 turns, while Qwen3.5-122B-A10B-NVFP4 has no issues finishing within ~50 turns or so. But why are you using Qwen3.6-27B-NVFP4 compared to the FP8 or full version? In my experience the Q8 of 27B is on par sometimes better than 122B. I am experiemnting witb higher quants for 122B to fit on my Strix Halo, but still, the difference honestly for my workflow is not that much. I just wish they released 3.6-122B version.
- nicman23 3mo agodense small models do not like quantization. i find 27b fp8 to be smarter albeit less knowledgable versus the 122B
- pettijohn 3mo agoSO MUCH THIS. I have Strix Halo with 128GB RAM and was a large and fast model like 122B A10B. Here's hoping!
- cmrdporcupine 3mo agoAbsolutely. There's a glaring gap in the space for something about the size of Nemotron Super or just under, but actually ... competent. The fantasy is a 100B or 80B model, but MoE and highly tuned for coding.