3 ms·
Try 27b, it's significantly smarter than 35b-a3b (although it is slower, it's not so bad with MTP).
by smcleod 4mo ago
Try 27b, it's significantly smarter than 35b-a3b (although it is slower, it's not so bad with MTP).
- ignoramous 4mo agoAt least according to gertlabs, Qwen3.6 27B outperforms every SoTA (closed) model at Kotlin: https://archive.vn/RYBCL https://archive.vn/RYBCL / https://gertlabs.com/rankings?mode=agentic_coding&language=kotlin https://gertlabs.com/rankings?mode=agentic_coding&language=k...
- iosjunkie 4mo agoInteresting. I wonder if there is opportunity to train a set of small model variants to excel at a certain stacks. Eg Qwen3.6-27B for Node + React or Qwen3.6-27B for Rust + TUI
- mft_ 4mo agoThis is always how I've imagined small/consumer-hardware models going in time. If I only ever code in Python, give me a model that does just that (plus some general CS, algorithms, structure, etc.) and does it super-fast and well. Make it small enough that if I need a Python back end and an HTML front end, another specific model can load alongside and collaborate on the front end. Or give me a pure shopping model that has a general understanding of products and product categories, and then will playwright/scrape/API into shopping sites to compare options and find me what I want. Etc.
- gertlabs 4mo agoQwen 3.6 27B is an anomalously strong all-around model for its size, but when we run our evaluations, we generate 10 coding submissions/language/model (110 total). So full discosure, the per-language per-model performances can be noisy (I do not think Qwen3.6 27B is better than Fable 5 in agentic workflows when writing Kotlin, given enough samples, although we do find some interesting anomalies that hold up under large sample sizes).
- bakies 4mo agoHmm, I just assumed bigger was better. How's it different?
- Lalabadie 4mo agoOff the top of my head since it seems to be the quick info you're looking for: IIRC, with these two, the 27B is a dense model, meaning it's all active at inference. Meanwhile, the 35B is a Mixture of Experts (MoE), so only part of its network (3B?) is active at any time.
- bakies 4mo agoThanks! Dense models have been slow on my compute, but I'll give it a try. If its not toooooo slow then it's fine I mostly fire and forget agents anyway. Edit: seems fast! I'll try it out some more, thanks again.
- smcleod 4mo ago35b-a3b is only 3b active parameters, it's a MoE.
- stymaar 4mo agoIt is, but it's way too slow on a Strix Halo due to its limited bandwidth. (I'm still sad that they didn't make a 122B-A10B version of it, as it's the kind of model that fits best on a Strix Halo, and for 3.5 it was comparable in performance to the dense 27B version).