12 ms·
Qwen's ~30B-class models are genuinely good enough for use if you can find a machine with enough memory bandwidth to run them at 30-90 tokens/second. It's been
by hadlock 4mo ago
Qwen's ~30B-class models are genuinely good enough for use if you can find a machine with enough memory bandwidth to run them at 30-90 tokens/second. It's been extremely telling that Qwen stopped releasing 120b class models. At some point in the next 10 years (maybe 3?) someone is going to release an Opus 4.5 class 256B model you can run locally. Right now our engineers use about $800/mo worth of opus tokens; at that rate the ROI for local LLM is ~10 months
- strictnein 4mo agoDidn't Qwen stop releasing their more powerful models because they're commercializing them?
- mswphd 4mo agoYes and no. Qwen 3.5 was released 3/2/2026. It includes models up to a 397B-A17B model https://huggingface.co/collections/Qwen/qwen35 https://huggingface.co/collections/Qwen/qwen35 A day afterwards, a high-up technical leader working on Qwen was let go https://techcrunch.com/2026/03/03/alibabas-qwen-tech-lead-steps-down-after-major-ai-push/ https://techcrunch.com/2026/03/03/alibabas-qwen-tech-lead-st... The more recent Qwen 3.6 was released on 4/16 https://huggingface.co/collections/Qwen/qwen36 https://huggingface.co/collections/Qwen/qwen36 This does not include any particularly large models. But the models it contains (Qwen3.6 27B and Qwen3.6 35B-A3B) are the local models people have been very excited about lately. So they didn't release any larger models, and the models people praise so much are from this most recent release.
- tyre 4mo agoIf they stop releasing their larger models because they want to monetize, would we expect them to release better small models that can outcompete those?
- mswphd 4mo agothere's pros and cons to it for them. Clearly, they get good branding (at least in enthusiast circles). Perhaps more important is they get community work on optimization. There have been significant performance uplifts on the Qwen3.6 models from the open-source community since they were launched (at a minimum, multi-token prediction is now working with them. It is almost a 2x token generation speedup) https://www.reddit.com/r/LocalLLM/comments/1ti9w4o/qwen3635ba3bmtp_on_an_rtx_3090_in_lm_studio_is/ https://www.reddit.com/r/LocalLLM/comments/1ti9w4o/qwen3635b...
- horsawlarway 4mo agoI want to echo this. I've been on claude's opus 4.5/6/7 for work for a couple months, and I finally got back to running Qwen A3B 35B... it's incredibly performant and quite capable on semi-reasonable local hardware. I get ~150 tokens/s on dual nvidia RTX 3090s and can fit the whole 300k context into gpu on a UD-Q4-K-XL quant gguf. Combined with Pi as a harness, and I'm surprised to find that it feels about as capable as claude did 8 months ago (their 3.x models). It's not Opus 4.5 levels yet, but it's good enough for a LOT of basic work. I actually downgraded my personal anthropic subscription because Qwen is absolutely fine for implementation work. I still let a better model write a plan, but then I can just switch over to Qwen to implement. I don't think we're 10 years away from opus 4.5 levels running on cheap consumer hardware. I think we're probably closer to 18 months away, and I suspect it'll be in the 30-60b range, not the 256b range. PC manufacturers also seem to be betting on local, with a LOT of focus on 64 to 128gb unified RAM machines.
- maxdo 4mo agoMajority of my agentic setup is pi / Claude code where every single Chinese models are not as good except commercial 1T models . Local is a pipe dream . If you can run it cheap occasionally why commercial companies can’t run it cheaper 24/7 and lower the costs ? The answer is simple. Use cases are more demanding and hence you need more from model not less . Sure if you task is to do a narrow labeling task on 1m records small optimized model is good . If you want to do complex things , it shifts with models advancements
- hparadiz 4mo agoThis sounds like something someone at IBM in 1986 would say trying to sell their mainframes. "PCs will never be a thing. No one's gonna want a computer." I'm seeing some impressive results from folks that can afford 10k+ GPUs right now. But those GPUs will all be hand me downs in 10 years. So pipe dream? Hmmm...... that's not how this industry works.
- tyre 4mo agoThose are not GPUs available on iPhones. Will we get there eventually? Maybe! Maybe we end up with GPU clusters built on the edge (e.g. cell towers) for offloading, maybe it’s never economical, maybe a different model architecture makes it simpler, who knows. But it doesn’t seem anywhere imminent with our current world state.