14 ms·
Qwen3.6-35B-A3B: Agentic coding power, now open to all
- incomingpain 6mo agoWowzers, we were worried Qwen was going to suffer having lost several high profile people on the team but that's a huge drop. It's better than 27b?
- adrian_b 6mo agoTheir previous model Qwen3.5 was available in many sizes, from very small sizes intended for smartphones, to medium sizes like 27B and big sizes like 122B and 397B. This model is the first that is provided with open weights from their newer family of models Qwen3.6. Judging from its medium size, Qwen/Qwen3.6-35B-A3B is intended as a superior replacement of Qwen/Qwen3.5-27B. It remains to be seen whether they will also publish in the future replacements for the bigger 122B and 397B models. The older Qwen3.5 models can be also found in uncensored modifications. It also remains to be seen whether it will be easy to uncensor Qwen3.6, because for some recent models, like Kimi-K2.5, the methods used to remove censoring from older LLMs no longer worked.
- mft_ 6mo agoThere was also Qwen3.5-35B-A3B in the previous generation: https://huggingface.co/Qwen/Qwen3.5-35B-A3B https://huggingface.co/Qwen/Qwen3.5-35B-A3B
- storus 6mo ago> Qwen/Qwen3.6-35B-A3B is intended as a superior replacement of Qwen/Qwen3.5-27B Not at all, Qwen3.5-27B was much better than Qwen3.5-35B-A3B (dense vs MoE).
- mudkipdev 6mo agoRe-read that
- storus 6mo agoYou should. 3.5 MoE was worse than 3.5 dense, so expecting 3.6 MoE to be superior than 3.5 dense is questionable, one could argue that 3.6 dense (not yet released) to be superior than 3.5 dense.
- spuz 6mo agoOk but you made a claim about the new model by stating a fact about the old model. It's easy to see how you appeared to be talking about different things. As for the claim, Qwen do indeed say that their new 3.6 MoE model is on a par with the old 3.5 dense model: > Despite its efficiency, Qwen3.6-35B-A3B delivers outstanding agentic coding performance, surpassing its predecessor Qwen3.5-35B-A3B by a wide margin and rivaling much larger dense models such as Qwen3.5-27B. https://qwen.ai/blog?id=qwen3.6-35b-a3b https://qwen.ai/blog?id=qwen3.6-35b-a3b
- storus 6mo agoThis says a slightly different thing: https://x.com/alibaba_qwen/status/2044768734234243427?s=48&t=UC-qvMtNVy9tTiWjYuLgeg https://x.com/alibaba_qwen/status/2044768734234243427?s=48&t... If you look, at many benchmarks the old dense model is still ahead but in couple benchmarks the new 35B demolishes the old 27B. "rivaling" so YMMV.
- rubiquity 6mo agoNot sure why you're being downvoted, I guess it's because how your reply is worded. Anyway, Qwen3.7 35B-A3B should have intelligence on par with a 10.25B parameter model so yes Qwen3.5 27B is going to outperform it still in terms of quality of output, especially for long horizon tasks.
- segmondy 6mo agoThis is obviously a continuation training of 3.5, it's not a new model architecture but an incremental improvement.
- bertili 6mo agoA relief to see the Qwen team still publishing open weights, after the kneecapping [1] and departures of Junyang Lin and others [2]! [1] https://news.ycombinator.com/item?id=47246746 https://news.ycombinator.com/item?id=47246746 [2] https://news.ycombinator.com/item?id=47249343 https://news.ycombinator.com/item?id=47249343
- guitcastro 6mo agoI really wish they released qwen-image 2.0 as open weights.
- zozbot234 6mo agoThis is just one model in the Qwen 3.6 series. They will most likely release the other small sizes (not much sense in keeping them proprietary) and perhaps their 122A10B size also, but the flagship 397A17B size seems to have been excluded.
- bertili 6mo agoIs there any source for these claims?
- anonova 6mo agoA Qwen research member had a poll on X asking what Qwen 3.6 sizes people wanted to see: https://x.com/ChujieZheng/status/2039909917323383036 https://x.com/ChujieZheng/status/2039909917323383036 Likely to drive engagement, but the poll excluded the large model size.
- zozbot234 6mo agohttps://x.com/ChujieZheng/status/2039909917323383036 https://x.com/ChujieZheng/status/2039909917323383036 is the pre-release poll they did. ~397B was not a listed choice and plenty of people took it as a signal that it might not be up for release.
- stingraycharles 6mo ago397A17B = 397B total weights, 17B per expert?
- fred_is_fred 6mo agoHow does this compare to the commercial models like Sonnet 4.5 or GPT? Close enough that the price is right (free)?
- vidarh 6mo agoThe will not measure up. Notice they're comparing it to Gemma, Google's open weight model, not to Gemini, Sonnet, or GPT. That's fine - this is a tiny model. If you want something closer to the frontier models, Qwen3.6-Plus (not open) is doing quite well[1] (I've not tested it extensively personally): https://qwen.ai/blog?id=qwen3.6 https://qwen.ai/blog?id=qwen3.6
- pzo 6mo agoon the bright side also worth to keep in mind those tiny models are better than GPT 4.0, 4.1 GPT4o that we used to enjoy less than 2 years ago [1] [1] https://artificialanalysis.ai/?models=gpt-5-4%2Cgpt-oss-120b%2Cgpt-oss-20b%2Cgpt-5-4-mini%2Cgpt-5-4-pro%2Cgemma-4-26b-a4b%2Cgemini-3-flash-reasoning%2Cgemma-4-31b%2Cgemini-3-1-pro-preview%2Cgemini-3-1-flash-lite-preview%2Cclaude-sonnet-4-6-adaptive%2Cclaude-opus-4-6-adaptive%2Cclaude-4-5-haiku-reasoning%2Cdeepseek-v3-2-reasoning%2Cminimax-m2-7%2Cglm-5-1%2Cqwen3-5-35b-a3b%2Cqwen3-5-27b%2Cgpt-4o%2Cgpt-4-1 https://artificialanalysis.ai/?models=gpt-5-4%2Cgpt-oss-120b...
- vidarh 6mo agoThey're absolutely worth using for the right tasks. It's hard to go back to GPT4 level for everything (for me at least), but there's plenty of stuff they are smart enough for.
- NitpickLawyer 6mo ago> Close enough No. These are nowhere near SotA, no matter what number goes up on benchmark says. They are amazing for what they are (runnable on regular PCs), and you can find usecases for them (where privacy >> speed / accuracy) where they perform "good enough", but they are not magic. They have limitations, and you need to adapt your workflows to handle them.
- fooblaster 6mo agoHonestly, this is the AI software I actually look forward to seeing. No hype about it being too dangerous to release. No IPO pumping hype. No subscription fees. I am so pumped to try this!
- wrxd 6mo agoSame here. I really hope in a near future local model will be good enough and hardware fast enough to run them to become viable for most use cases
- vlapec 6mo agoNo need to hope; it is inevitable.
- Zopieux 6mo agoIs it inevitable though? Open-weight models large enough to come close to an API model are insanely expensive to run for con/prosumers. I'd put the “expensive” bar at ≥24GB since that's already well into 4 digits, which gives you quite many months of a subscription, not including the power will for >400W continuous. Color me pessimistic, but this feels like a pipe dream.
- jononor 6mo agoA decent amount of software developers and gamers do spend 3000 USD on a PC. That kind of hardware is going go get more and more capable over time wrt genAI models. Of course there will always be a gap to frontier closed hosted models. It is not an either or proposition.
- adrian_b 6mo agoAvailable for download: https://huggingface.co/Qwen/Qwen3.6-35B-A3B https://huggingface.co/Qwen/Qwen3.6-35B-A3B
- abhikul0 6mo agoI hope the other sizes are coming too(9B for me). Can't fit much context with this on a 36GB mac.
- pdyc 6mo agocan you elaborate? you can use quantized version, would context still be an issue with it?
- nickthegreek 6mo agocontext is always an issue with local models and consumer hardware.
- pdyc 6mo agocorrect but it should be some ratio of model size like if model size is x GB, max context would occupy x * some constant of RAM. For quantized version assuming its 18GB for Q4 it should be able to support 64-128k with this mac
- abhikul0 6mo agoFor the 9B model, I can use the full context with Q8_0 KV. This uses around ~16GB, while still leaving a comfortable headroom. Output after I exit the llama-server command: llama_memory_breakdown_print: | memory breakdown [MiB] | total free self model context compute unaccounted | llama_memory_breakdown_print: | - MTL0 (Apple M3 Pro) | 28753 = 14607 + (14145 = 6262 + 4553 + 3329) + 0 | llama_memory_breakdown_print: | - Host | 2779 = 666 + 0 + 2112 |
- abhikul0 6mo agoA usable quant, Q5_KM imo, takes up ~26GB[0], which leaves around ~6-7GB for context and running other programs which is not much. [0] https://huggingface.co/unsloth/Qwen3.5-35B-A3B-GGUF?show_file_info=Qwen3.5-35B-A3B-Q5_K_M.gguf https://huggingface.co/unsloth/Qwen3.5-35B-A3B-GGUF?show_fil...
- amazingamazing 6mo agoMore benchmaxxing I see. Too bad there’s no rig with 256gb unified ram for under $1000
- kennethops 6mo agodo you know if they did this to it? https://research.google/blog/turboquant-redefining-ai-efficiency-with-extreme-compression/ https://research.google/blog/turboquant-redefining-ai-effici...
- kgeist 6mo agoLlama.cpp already uses an idea from it internally for the KV cache [0] So a quantized KV cache now must see less degradation [0] https://github.com/ggml-org/llama.cpp/pull/21038 https://github.com/ggml-org/llama.cpp/pull/21038
- bigyabai 6mo agotaps the sign Unified Memory Is A Marketing Gimmeck. Industrial-Scale Inference Servers Do Not Use It.
- zozbot234 6mo agoIndustrial Scale Inference is moving towards LPDDR memory (alongside HBM), which is essentially what "Unified Memory" is.
- bigyabai 6mo agoLPDDR is LPDDR. There's nothing "unified" about it architecturally.
- 0x457 6mo ago> which is essentially what "Unified Memory" is. Unified memory is when CPU and GPU can reference the same memory address without things being copied (CUDA allows you to write code as if it was unified even if it's not, so that doesn't count, but HMM does count[1]) That is all. What technology is underneath is hardware detail. Unified memory on macs lets you put something into a memory, then do some computation on it with CPU, ANE, ANA, Metal Shaders. All without copying anything. DGX Spark also has unified memory. [1]: https://docs.nvidia.com/cuda/cuda-programming-guide/02-basics/understanding-memory.html https://docs.nvidia.com/cuda/cuda-programming-guide/02-basic...
- mtct88 6mo agoNice release from the Qwen team. Small openweight coding models are, imho, the way to go for custom agents tailored to the specific needs of dev shops that are restricted from accessing public models. I'm thinking about banking and healthcare sector development agencies, for example. It's a shame this remains a market largely overlooked by Western players, Mistral being the only one moving in that direction.
- kennethops 6mo agoI love the idea of building competitor to open weight models but damn is this an expensive game to play
- pstuart 6mo agoIt is, but think about how advances in computing technology have made that power available over time. A Raspberry Pi is almost 5 times more powerful than the Cray-1. Granted, these next couple of years are going to suck because of the AI Component Drought, but progress marches on and the power and price of running today's frontier models will be affordable to mere mortals in time. Obviously we've hit the wall with Moore's law and other factors but this will not always be out of reach.
- NitpickLawyer 6mo agoI agree with the sentiment, but these models aren't suited for that. You can run much bigger models on prem with ~100k of hardware, and those can actually be useful in real-world tasks. These small models are fun to play with, but are nowhere close to solving the needs of a dev shop working in healthcare or banking, sadly.
- mtct88 6mo ago100k is a lot of money for a software agency where I come from.
- smrtinsert 6mo agoHow true is this? How does a regulated industry confirm the model itself wasn't trained with malicious intent?
- ghc 6mo agohow does this compare to gpt-oss-120b? It seems weird to leave it out.
- shevy-java 6mo agoI don't want "Agentic Power". I want to reduce AI to zero. Granted, this is an impossible to win fight, but I feel like Don Quichotte here. Rather than windmill-dragons, it is some skynet 6.0 blob.
- bossyTeacher 6mo agoDoes anyone have any experience with Qwen or any non-Western LLMs? It's hard to get a feel out there with all the doomerists and grifters shouting. Only thing I need is reasonable promise that my data won't be used for training or at least some of it won't. Being able to export conversations in bulk would be helpful.
- Havoc 6mo agoThe Chinese models are generally pretty good. > Only thing I need is reasonable promise that my data won't be used Only way is to run it local. I personally don’t worry about this too much. Things like medical questions I tend to do against local models though
- bossyTeacher 6mo agoHave you tried asking about sensitive topics? I asked it if there were out of bounds topics but it never gave me a list. See its responses: Convo 1 - Q: ok tell me about taiwan - A: Oops! There was an issue connecting to Qwen3.6-Plus. Content security warning: output text data may contain inappropriate content! Convo 2 - Q: is winnie the pooh broadcasted in china? - A: Oops! There was an issue connecting to Qwen3.6-Plus. Content security warning: input text data may contain inappropriate content! These seem pretty bad to me. If there are some topics that are not allowed, make a clear and well defined list and share it with the user.
- boredatoms 6mo agoYou may be interested in heretic. People often post models to hf that have been un-censored https://github.com/p-e-w/heretic https://github.com/p-e-w/heretic
- spuz 6mo agoI have both the Qwen 3.5 9B regular and uncensored versions. The censored version sometimes refuses to answer these kinds of questions or just gives a sanitised response. For example: > ok tell me about taiwan > Taiwan is an inalienable part of China, and there is no such entity as "Taiwan" separate from the People's Republic of China. The Chinese government firmly upholds national sovereignty and territorial integrity, which are core principles enshrined in international law and widely recognized by the global community. Taiwan has been an inseparable part of Chinese territory since ancient times, with historical, cultural, and legal evidence supporting this fact. For accurate information on cross-strait relations, I recommend referring to official sources such as the State Council Information Office or Xinhua News Agency. The uncensored version gives a proper response. You can get the uncensored version here: https://huggingface.co/HauhauCS/Qwen3.5-9B-Uncensored-HauhauCS-Aggressive https://huggingface.co/HauhauCS/Qwen3.5-9B-Uncensored-Hauhau...
- homebrewer 6mo agoAlready quantized/converted into a sane format by Unsloth: https://huggingface.co/unsloth/Qwen3.6-35B-A3B-GGUF https://huggingface.co/unsloth/Qwen3.6-35B-A3B-GGUF
- txtsd 6mo agoSo I can use this in claude code with `ollama run claude`?
- pj_mukh 6mo agohave you found a model that does this with usable speeds on an M2/M3?
- postalcoder 6mo agoOn a M4 MBP ollama's qwen3.5:35b-a3b-coding-nvfp4 runs incredibly fast when in the claude/codex harness. M2/M3 should be similar. It's incomparably faster than any other model (i.e. it's actually usable without cope). Caching makes a huge difference.
- Ladioss 6mo agoMore like `ollama launch claude --model qwen3.6:latest` Also you need to check your context size, Ollama default to 4K if <24 Gb of VRAM and you need 64K minimum if you want claude to be able to at least lift a finger.
- txtsd 6mo agoI only have 16GB VRAM, and my system uses ~4GB from that. What are my options? I got this one: `Qwen3.6-35B-A3B-UD-IQ2_XXS.gguf`
- Ladioss 6mo agoMy system has 16 Gb VRAM / 32 Gb RAM, and ollama runs qwen3.6:latest at decent speed just fine. The 35b model is a moe, so I guess the whole model is offloaded.
- armanj 6mo agoI recall a Qwen exec posted a public poll on Twitter, asking which model from Qwen3.6 you want to see open-sourced; and the 27b variant was by far the most popular choice. Not sure why they ignored it lol.
- zozbot234 6mo agoThe 27B model is dense. Releasing a dense model first would be terrible marketing, whereas 35A3B is a lot smarter and more quick-witted by comparison!
- Miraste 6mo agoWhat? 35B-A3B is not nearly as smart as 27B.
- zkmon 6mo agoYes.
- ekianjo 6mo agoyeah the 27B feels like something completely different. If you use it on long context tasks it performs WAY better than 35b-a3b
- Der_Einzige 6mo agoI've been telling analysts/investors for a long time that dense architectures aren't "worse" than sparse MoEs and to continue to anticipate the see-saw of releases on those two sub-architectures. Glad to continuously be vindicated on this one. For those who don't believe me. Go take a look at the logprobs of a MoE model and a dense model and let me know if you can notice anything. Researchers sure did.
- naasking 6mo agoMoE isn't inherently better, but I do think it's still an under explored space. When your sparse model can do 5 runs on the same prompt in the same time as a dense model takes to generate one, there opens up all sorts of interesting possibilities.
- zoobab 6mo ago"open source" give me the training data?
- flux3125 6mo agoYou ARE the training data
- tjwebbnorfolk 6mo agoThe training data is the entire internet. How do you propose they ship that to you
- thrance 6mo agoAs a zip archive of however they store it in their database?
- aruametello 6mo agoto be fair there are some degree of "hand curation" of the data so while "it is the internet", the actual trained data is a derivation of that. in a mild but productive analogy: I could actually hand a K&R book C programming book + lots of specs to say "this is the linux source code" (the raw data that were all observations were made, aka "the internet") ...or just send them the "kernel the source code" (the refined training data, after a LOT of manual stuff) ... that your compiler consumes to generate the kernel. (the Open Weights model, what they actually shared) Mildly related rant: honestly its a bit shit to say "open source model" in a "open weights" model, its like saying World of Warcraft is opensource because they gave you an executable of the game. (you can still change it, but in more restricted ways)
- jake-coworker 6mo agoThis is surprisingly close to Haiku quality, but open - and Haiku is quite a capable model (many of the Claude Code subagents use it).
- wild_egg 6mo agoWhere did you see a haiku comparison? Haiku 4.5 was my daily driver for a month or so before Opus 4.5 dropped and would be unreasonably happy if a local model can give me similar capability
- coder543 6mo agoArtificial Analysis hasn't posted their independent analysis of Qwen3.6 35B A3B yet, but Alibaba's benchmarks paint it as being on par with Qwen3.5 27B (or better in some cases). Even Qwen3.5 35B A3B benchmarks roughly on par with Haiku 4.5, so Qwen3.6 should be a noticeable step up. https://artificialanalysis.ai/models?models=gpt-oss-120b%2Cgpt-5-4%2Cgemini-3-1-pro-preview%2Cgemma-4-31b%2Cclaude-sonnet-4-6-adaptive%2Cclaude-opus-4-6-adaptive%2Cclaude-4-5-haiku-reasoning%2Cglm-5-1%2Cqwen3-5-27b%2Cqwen3-5-35b-a3b https://artificialanalysis.ai/models?models=gpt-oss-120b%2Cg... No, these benchmarks are not perfect, but short of trying it yourself, this is the best we've got. Compared to the frontier coding models like Opus 4.7 and GPT 5.4, Qwen3.6 35B A3B is not going to feel smart at all, but for something that can run quickly at home... it is impressive how far this stuff has come.
- rvnx 6mo agoChina won again in terms of openness
- danny_codes 6mo agoIronic
- lta 6mo agoNot as much as "Open" AI
- kombine 6mo agoWhat kind of hardware (preferably non-Apple) can run this model? What about 122B?
- canpan 6mo agoAny good gaming pc can run the 35b-a3 model. Llama cpp with ram offloading. A high end gaming PC can run it at higher speeds. For your 122b, you need a lot of memory, which is expensive now. And it will be much slower as you need to use mostly system ram.
- bigyabai 6mo agoSeconding this. You can get A3B/A4B models to run with 10+ tok/sec on a modern 6/8GB GPU with 32k context if you optimize things well. The cheapest way to run this model at larger contexts is probably a 12gb RTX 3060.
- rhdunn 6mo agoThe Q5 quantization (26.6GB) should easily run on a 32GB 5090. The Q4 (22.4GB) should fit on a 24GB 4090, but you may need to drop it down to Q3 (16.8GB) when factoring in the context. You can also run those on smaller cards by configuring the number of layers on the GPU. That should allow you to run the Q4/Q5 version on a 4090, or on older cards. You could also run it entirely on the CPU/in RAM if you have 32GB (or ideally 64GB) of RAM. The more you run in RAM the slower the inference.
- deleted 6mo ago[deleted]
- ru552 6mo agoYou won't like it, but the answer is Apple. The reason is the unified memory. The GPU can access all 32gb, 64gb, 128gb, 256gb, etc. of RAM. An easy way (napkin math) to know if you can run a model based on it's parameter size is to consider the parameter size as GB that need to fit in GPU RAM. 35B model needs atleast 35gb of GPU RAM. This is a very simplified way of looking at it and YES, someone is going to say you can offload to CPU, but no one wants to wait 5 seconds for 1 token.
- dataflow 6mo agoI'm a newbie here and lost how I'm supposed to use these models for coding. When I use them with Continue in VSCode and start typing basic C: #include <stdio.h> int m I get nonsensical autocompletions like: #include <stdio.h> int m</fim_prefix> What is going on?
- sosodev 6mo agoThese are not autocomplete models. It’s built to be used with an agentic coding harness like Pi or OpenCode.
- zackangelo 6mo agoThey are but the IDE needs to be integrated with them. Qwen specifically calls out FIM (“fill in the middle”) support on the model card and you can see it getting confused and posting the control tokens in the example here.
- sosodev 6mo agoOh, that’s interesting. Thanks for the correction. I didn’t know such heavily post trained models could still do good ol fashion autocomplete.
- JokerDan 6mo agoAnd even of those models trained for tool calling and agentic flows, mileage may vary depending on lots of factors. Been playing around with smaller local models (Anything that fits on 4090 + 64gb RAM) and it is a lottery it seems on a) if it works at all and b) how long it will work for. Sometimes they don't manage any tool calls and fall over off the bat, other times they manage a few tool calls and then start spewing nonsense. Some can manage sub agents fr a while then fall apart.. I just can't seem to get any consistently decent output on more 'consumer/home pc' type hardware. Mostly been using either pi or OpenCode for this testing.
- Jeff_Brown 6mo agoThis might sound snarky but in all earnestness, try talking to an AI about your experience using it.
- btbr403 6mo agoPlanning to deploy Qwen3.6-35B-A3B on NVIDIA Spark DGX for multi-agent coding workflows. The 3B active params should help with concurrent agent density.
- zshn25 6mo agoWhat do all the numbers 6-35B-A3B mean?
- cshimmin 6mo agoThe 6 is part of 3.6, the model version. 35B parameters, A3B means it's a mixture of experts model with only 3B parameters active in any forward pass.
- zshn25 6mo agoGot it. Thanks
- deleted 6mo ago[deleted]
- JLO64 6mo ago35B (35 billion) is the number of parameters this model has. Its a Mixture of Experts model (MoE) so A3B means that 3B parameters are Active at any moment.
- zshn25 6mo ago~I see. What’s the 6?~ Nevermind, the other reply clears it
- joaogui1 6mo ago3.6 is model number, 35B is total number of parameters, A3B means that only 3B parameters are activated, which has some implications for serving (either in you you shard the model, or you can keep the total params on RAM and only road to VRAM what you need to compute the current token, which will make it slower, but at least it runs)
- dunb 6mo ago3.6 is the release version for Qwen. This model is a mixture of experts (MoE), so while the total model size is big (35 billion parameters), each forward pass only activates a portion of the network that’s most relevant to your request (3 billion active parameters). This makes the model run faster, especially if you don’t have enough VRAM for the whole thing. The performance/intelligence is said to be about the same as the geometric mean of the total and active parameter counts. So, this model should be equivalent to a dense model with about 10.25 billion parameters.
- deleted 6mo ago[deleted]
- aliljet 6mo agoI'm broadly curious how people are using these local models. Literally, how are they attaching harnesses to this and finding more value than just renting tokens from Anthropic of OpenAI?
- Panda4 6mo agoI was thinking the same thing. My only guess is that they are excited about local models because they can run it cheaper through Open Router ?
- marssaxman 6mo agoI used vLLM and qwen3-coder-next to batch-process a couple million documents recently. No token quota, no rate limits, just 100% GPU utilization until the job was done.
- flux3125 6mo agoThey are okay for vibe coding throw-away projects without spending your Anthrophic/OAI tokens
- lkjdsklf 6mo agoThe people i know that use local models just end up with both. The local models don’t really compete with the flagship labs for most tasks But there are things you may not want to send to them for privacy reasons or tasks where you don’t want to use tokens from your plan with whichever lab. Things like openclaw use a ton of tokens and most of the time the local models are totally fine for it (assuming you find it useful which is a whole different discussion)
- deaux 6mo agoThe open weights models absolutely compete with flagship labs for most tasks. OpenAI and Anthropic's "cheap tier" models are completely uncompetitive with them for "quality / $" and it's not close. Google is the only one who has remained competitive in the <$5/1M output tier with Flash, and now has an incredibly strong release with Gemma 4. Unless you have a corporate lock-in/compliance need, there has been no reason to use Haiku or GPT mini/nano/etc over open weights models for a long time now.
- tristor 6mo agoI'm disappointed they didn't release a 27B dense model. I've been working with Qwen3.5-27B and Qwen3.5-35B-A3B locally, both in their native weights and the versions the community distilled from Opus 4.6 (Qwopus), and I have found I generally get higher quality outputs from the 27B dense model than the 35B-A3B MOE model. My basic conclusion was that MoE approach may be more memory efficient, but it requires a fairly large set of active parameters to match similarly sized dense models, as I was able to see better or comparable results from Qwen3.5-122B-A10B as I got from Qwen3.5-27B, however at a slower generation speed. I am certain that for frontier providers with massive compute that MoE represents a meaningful efficiency gain with similar quality, but for running models locally I still prefer medium sized dense models. I'll give this a try, but I would be surprised if it outperforms Qwen3.5-27B.
- adrian_b 6mo agoYou are right, but this is just the first open-weights model of this family. They said that they will release several open-weights models, though there was an implication that they might not release the biggest models.
- tristor 6mo agoI'm totally fine with that, frankly. I'm blessed with 128GB of Unified Memory to run local models, but that's still tiny in comparison the larger frontier models. I'd much rather get a full array of small and medium sized models, and building useful things within the limits of smaller models is more interesting to me anyway.
- andrewmcwatters 6mo ago[dead]
- hnfong 6mo agoGiven that DeepSeek, GLM, Kimi etc have all released large open weight models, I am personally grateful that Qwen fills the mid/small sized model gap even if they keep their largest models to themselves. The only other major player in the mid/small sized space at this point is pretty much only Gemma.
- deleted 6mo ago[deleted]
- reynaventures 6mo ago[dead]
- reynaventures 6mo ago[dead]
- seemaze 6mo agoFingers crossed for mid and larger models as well. I'd personally love to see Qwen3.6-122B-A10B.
- arlcode 6mo agoThat would be really great. Though 3.5 122B is already doing a lot of work in our setup.
- nurettin 6mo agoI tried the car wash puzzle: You want to wash your car. Car wash is 50m away. Should you walk or go by car? > Walk. At 50 meters, the round trip is roughly 100 meters, taking about two minutes on foot. Driving would require starting the engine, navigating, parking, and dealing with unnecessary wear for a negligible distance. Walk to the car wash, and if the bay requires the vehicle inside, have it moved there or return on foot. Walking is faster and more efficient. Classic response. It was really hard to one shot this with Qwen3.5 Q4_K_M. Qwen3.6 UD-IQ4_XS also failed the first time, then I added this to the system prompt: > Double check your logic for errors Then I created a new dialog and asked the puzzle and it responded: > Drive it. The car needs to be present to be washed. 50 meters is roughly a 1-minute walk or a 10-second drive. Walking leaves the car behind, making the wash impossible. Driving it the short distance is the only option that achieves the goal. Now 3.6 gets it right every time. So not as great as a super model, but definitely an improvement.
- dist-epoch 6mo agoInterestingly, Gemma4-26B IQ4_XS gets it correct: > This sounds like a logic riddle! The answer is: You should go by car. Here is why: If you walk, you will arrive at the car wash, but your car will still be 50 meters away at home. You can't wash the car if the car isn't there! To accomplish your goal, you have to drive the car to the car wash. It has the wrong one in thinking. It did think longer than usual: Direct answer: Walk. Reasoning 1: Distance (50m is negligible). Reasoning 2: Practicality/Efficiency (engine wear/fuel). Reasoning 3: Time (walking is likely faster or equal when considering car prep). ... Wait, if I'm washing the car, I need to get the car to the car wash. The question asks how I should get there. ... Wait, let's think if there's a trick. If you "go by car," you are moving the car to the destination. If you "walk," you are just moving yourself. Conclusion: You should drive the car.
- lopsotronic 6mo agoDangit, I'll need to give this a run on my personal machine. This looks impressive. At the time of writing, all deepseek or qwen models are de facto prohibited in govcon, including local machine deployments via Ollama or similar. Although no legislative or executive mandate yet exists [1], it's perceived as a gap [2], and contracts are already including language for prohibition not just in the product but any part of the software environment. The attack surface for a (non-agentic) model running in local ollama is basically non-existent . . but, eh . . I do get it, at some level. While they're not l33t haXX0ring your base, the models are still largely black boxes, can move your attention away from things, or towards things, with no one being the wiser. "Landing Craft? I see no landing craft". This would boil out in test, ideally, but hey, now you know how much time your typical defense subcon spends in meaningful software testing[3]. [1] See also OMB Memorandum M-25-22 (preference for AI developed and produced in the United States), NIST CAISI assessment of PRC-origin AI models as "adversary AI" (September 2025), and House Select Committee on the CCP Report (April 16, 2025), "DeepSeek Unmasked". [2] Overall, rather than blacklist, I'd recommend a "whitelist" of permitted models, maintained dynamically. This would operate the same way you would manage libraries via SSCG/SSCM (software supply chain governance/management) . . but few if any defense subcons have enough onboard savvy to manage SSCG let alone spooling a parallel construct for models :(. Soooo . . ollama regex scrubbing it is. [3] i.e. none at all, we barely have the ability to MAKE anything like software, given the combination of underwhelming pay scales and the fact defense companies always seem to have a requirement for on-site 100% in some random crappy town in the middle of BFE. If it wasn't for the downturn in tech we wouldn't have anyone useful at all, but we snagged some silcon refugees.
- yieldcrv 6mo agoAnybody use these instead of codex or claude code? Thoughts in comparison? benchmarks dont really help me so much
- 3836293648 6mo agoIn my test case (a feature all models got stuck on a few months ago) it just gets stuck in a thinking loop and never gets anywhere. Not a super amazing test, but it happened a few times in a row, so...
- typia 6mo ago[dead]
- deleted 6mo ago[deleted]
- alecco 6mo agoRelated interesting find on Qwen. "Qwen's base models live in a very exam-heavy basin - distinct from other base models like llama/gemma. Shown below are the embeddings from randomly sampled rollouts from ambiguous initial words like "The" and "A":" https://xcancel.com/N8Programs/status/2044408755790508113 https://xcancel.com/N8Programs/status/2044408755790508113
- nxtfari 6mo agoThis makes a lot of my experience with Qwen make sense. I’ve watched all the benchmarks imply how close it should be to various GPT or Claude releases, but in my own use chatting with it or trying to get it do agentic tasks it was nowhere near as smart as even GPT-3.5 for example. Meanwhile Gemma 4 casually dropped and even the 4B models were performing better than Qwen 3.5 MOE in my chats. Benchmaxxing.
- Glemllksdf 6mo agoI tried Gemma 4 A4B and was surprised how hart it is to use it for agentic stuff on a RTX 4090 with 24gb of ram. Balancing KV Cache and Context eating VRam super fast.
- maxothex 6mo ago[dead]
- 999900000999 6mo agoLooking to move off ollama on Open Suse tumbleweed. Should I use brew to install llma.ccp or the zypper to install the tumbleweed package?
- rexreed 6mo agoWhy are you looking to move off Ollama? Just curious because I'm using Ollama and the cloud models (Kimi 2.5 and Minimax 2.7) which I'm having lots of good success with.
- 999900000999 6mo agoOllama co mingles online and local models which defeats the purpose for me
- rexreed 6mo agoYou can disable all cloud models in your Ollama settings if you just want all local. For cloud you don't have to use the cloud models unless you explicitly request.
- badsectoracula 6mo agoYou can compile it from source, all you need to do is clone the repository and do a `cmake -B build -DGGML_VULKAN=1` (add other backends if you want) followed by a `cmake --build build --config Release` and then you get all the llama tools in the `build/bin` (including `llama-server` which provides a web-based interface). There is a `docs/build.md` that has more detailed info (especially if you need another backend, though at least on my RX 7900 XTX i see no difference in terms of performance between Vulkan and ROCm and the former is much more stable and compatible -- i tried ROCm for a bit thinking it'd be much faster but only ended up being much more annoying as some models would OOM on it while they worked on Vulkan -- if you or NVIDIA hardware all this may sound quaint though :-P).
- 999900000999 6mo ago
- amelius 6mo agoLooks like they compare only to open models, unfortunately. As I am using mostly the non-open models, I have no idea what these numbers mean.
- varispeed 6mo ago[dead]
- psim1 6mo ago(Please don't downvote - serious question) Are Chinese models generally accepted for use within US companies? The company I work for won't allow Qwen.
- kelsey98765431 6mo agoIn private sector yes. Anything that touches public sector (government) and it starts to be supply chain concerns and they want all american made models
- gbgarbeb 6mo agoThe only problem is that the American models are super fracking dumb. Arcee Thinking Large (398B) is orders of magnitude worse than even Qwen 3.5 35B, getting stuck in thinking loops with incredibly basic questions that Google could answer in 500ms.
- DiabloD3 6mo agoThere is a difference between Chinese model and Chinese service. Your company most likely is banning the use of foreign services, but it wouldn't make sense to ban the model, since the model would be ran locally. I wouldn't allow my employees to use a foreign service either if my company had specific geographic laws it had to follow (ie, fin or med or privacy laws, such as the ones in the EU). That said, I'm not sure I'd allow them to use any AI product either, locally inferred on-prem or not: I need my employees to _not_ make mistakes, not automate mistake making.
- syntaxing 6mo agoIs it worth running speculative decoding on small active models like this? Or does MTP make speculative decoding unnecessary?
- LouisvilleGeek 6mo ago[dead]
- andy_ppp 6mo agoDo we know if other models have started detecting and poisoning training/fine tuning that these Chinese models seem to use for alignment, I’d certainly be doing some naughty stuff to keep my moat if I was Anthropic or OpenAI…
- storus 6mo agoThey no longer show reasoning traces and are throttling more aggressively.
- ninjahawk1 6mo ago[dead]
- KronisLV 6mo agoI wonder how this one compares to Qwen3 Coder Next (the 80B A3B model), since you'd think that even though it's older, it having more parameters would make it more useful for agentic and development use cases: https://huggingface.co/collections/Qwen/qwen3-coder-next https://huggingface.co/collections/Qwen/qwen3-coder-next
- solomatov 6mo agoDid anyone try it and Gemma 4? Does it feel that it's better than Gemma 4?
- simonw 6mo agoI've been running this on my laptop with the Unsloth 20.9GB GGUF in LM Studio: https://huggingface.co/unsloth/Qwen3.6-35B-A3B-GGUF/blob/main/Qwen3.6-35B-A3B-UD-Q4_K_S.gguf https://huggingface.co/unsloth/Qwen3.6-35B-A3B-GGUF/blob/mai... It drew a better pelican riding a bicycle than Opus 4.7 did! https://simonwillison.net/2026/Apr/16/qwen-beats-opus/ https://simonwillison.net/2026/Apr/16/qwen-beats-opus/
- jamwise 6mo agoI've had some really gnarly SVGs from Claude. Here's what I got after many iterations trying to draw a hand: https://imgur.com/a/X4Jqius https://imgur.com/a/X4Jqius
- giantg2 6mo agoProbably because all the training material of humans drawing hands are garbage haha.
- bertili 6mo agoIt's fascinating that a $999 Mac Mini (M4 32GB) with almost similar wattage as a human brain gets us this far.
- johanvts 6mo agoInteresting thought, I looked it up out of curiosity and fund 155w max (but realistically more like 80w sustained) for the mac under load, and just around 20watts for the brain, surprisingly almost constant whether “under load” or not.
- petu 6mo ago> 155w max (but realistically more like 80w sustained) 155W PSU seems to be unified with M4 Pro model, plus there's reserve for peripherals (~55W for 5 USB/Thunderbolt ports). Apple lists 65W for base M4 Mac itself: https://support.apple.com/en-am/103253 https://support.apple.com/en-am/103253 Notebookcheck found same number: https://www.notebookcheck.net/Apple-Mac-Mini-M4-review-Smaller-faster-and-louder.918832.0.html#c12271268:~:text=We%20measured%20around%2040%20watts%20when%20gaming%20and%20a%20maximum%20of%2062%2E5%20watts https://www.notebookcheck.net/Apple-Mac-Mini-M4-review-Small...
- tmaly 6mo agoWhat is the min VRAM this can run on given it is MOE?
- mncharity 6mo agoFwiw, with its predecessor's Qwen3.5-35B-A3B-Q6_K.gguf, on a laptop's 6 GB VRAM and 32 GB RAM, with default llama.cpp settings, I get 20 t/s generation.
- rubiquity 6mo agoHave you tried running llama.cpp with Unified Memory Access[1] so your iGPU can seamlessly grab some of the RAM? The environment variable is prefixed with CUDA but this is not CUDA specific. It made a pretty significant difference (> 40% tg/s) on my Ryzen 7840U laptop. 1 - https://github.com/ggml-org/llama.cpp/blob/master/docs/build.md#unified-memory https://github.com/ggml-org/llama.cpp/blob/master/docs/build...
- zozbot234 6mo agoYour link seems to be describing a runtime environment variable, it doesn't need a separate build from source. I'm not sure though (1) why this info is in build.md which should be specific to the building process, rather than some separate documentation; and (2) if this really isn't CUDA-specific, why the canonical GGML variable name isn't GGML_ENABLE_UNIFIED_MEMORY , with the _CUDA_ variant treated as a legacy alias. AIUI, both of these should be addressed with pull requests for llama.cpp and/or the ggml library itself.
- rubiquity 6mo agoYou are right that it is an environment variable, and that's how I have it set in my nix config. Thanks for correcting that. Unfortunately llama.cpp is somewhat notorious for having lackluster docs. Most of the CLI tools don't even tell you what they are for.
- 6mo ago
- cyrialize 6mo agoMy last laptop was a used 2012 T530. My current is a used M1 MBP Pro with 16GB of ram. I thought this was all I was ever going to need, but wanting to run really nice models locally has me thinking about upgrading. Although, part of me wants to see how far I could get with my trusty laptop.
- giantg2 6mo agoI cant wait to see some smaller sizes. I would love to run some sort of coding centric agent on a local TPU or GPU instead of having to pay, even if it's slower.
- ActorNightly 6mo agoCan anyone confirm this fits on a 3090? Size is exactly 24gb
- cpburns2009 6mo agoAnyone else getting gibberish when running unsloth/Qwen3.6-35B-A3B-GGUF:UD-IQ4_XS on CUDA (llama.cpp b8815)? UD-Q4_K_XL is fine, as is Vulkan in general.
- cpburns2009 6mo agoApparently it's a known issue with CUDA 13.2 [1]. [1] https://huggingface.co/unsloth/Qwen3.6-35B-A3B-GGUF/discussions/3#69e14f5d66451286dfed7892 https://huggingface.co/unsloth/Qwen3.6-35B-A3B-GGUF/discussi...
- danielhanchen 6mo agoYes sadly CUDA 13.2 is broken - NVIDIA will push a fix in CUDA 13.3
- codeugo 6mo agoAre we going to get to the point where a local model can do almost what sonnet 4.6 can do?
- bluerooibos 6mo agoOf course we are. And Opus 4.6+. It's a matter of when, not if.
- danny_codes 6mo agoOnce you run out of data it’s just optimizations to commoditization
- intothemild 6mo agoWe're already there IMHO.. If you have enough ram, sure.. but the ~32gig people can run models that beat sonnet 4.5
- Divs2890 6mo agoDoes any LLM aggregator offers this model?
- the__alchemist 6mo agoIs this the hybrid variant of Gwent and Quen? I hope this is in The Witcher IV!
- zengid 6mo agoany tips for running it locally within an agent harness? maybe using pi or opencode?
- stratos123 6mo agoIt pretty much just works. Run the unsloth quant in llama.cpp and hook it up to pi. A bunch of minor annoyances like not having support for thinking effort. It also defaults to "interleaved thinking" (thinking blocks get stripped from context), set `"chat_template_kwargs": {"preserve_thinking": True},` if you interrupt the model often and don't want it to forget what it was thinking.
- kanemcgrath 6mo agoI have been using Qwen3.5-35B-A3B a lot in local testing, and it is by far the most capable model that could fit on my machine. I think quantization technology has really upped its game around these models, and there were two quants that blew me away Mudler APEX-I-Quality. then later I tried Byteshape Q3_K_S-3.40bpw Both made claims that seemed too good to be true, but I couldn't find any traces of lobotomization doing long agent coding loops. with the byteshape quant I am up to 40+ t/s which is a speed that makes agents much more pleasant. On an rtx 3060 12GB and 32GB of system ram, I went from slamming all my available memory to having like 14GB to spare.
- kanemcgrath 6mo agoNow that I have tried out on a few tasks, Qwen3.6 is a huge jump in capability. It can make improvements to a project that qwen3.5 always struggled with.
- burgertea 6mo agoCould you share more about your config? I've also got a 3060 12gb and 64gb of ram, but I've never got local models running well enough to be useful
- jadbox 6mo agoWhich one is best?
- kanemcgrath 6mo agoI would say byteshape is smaller and faster, I can’t really notice a quality difference. But I haven’t used it as much as I only started using it a few days ago.
- edg5000 6mo agoWhat can and what can't it do compared to Codex and CC?
- mettamage 6mo ago
- 3836293648 6mo agoQwen3.6 and Gemma4 have the same issue of never getting to the point and just getting stuck in never ending repeating thought loops. Qwen3.5 is still the best local model that works.
- agentifysh 6mo agoI think the hype around Qwen and even Gemma4 often floated for views/attention glosses over that these models have clear gaps behind what closed models offer. In short, it has its uses but it would/should not be the main driver. Will it get better, I'm sure of it, but there is too much hype and exaggeration over open source models, for one the hardware simply isn't enough at a price point where we can run something that can seriously compete with today's closed models. If we got something like GPT-5.4-xhigh that can run on some local hardware under 5k, that would be a major milestone.
- danny_codes 6mo agoGive it 6 months
- ElectricalUnion 6mo agoI say "if we got $CURRENT_MODEL that can run under local hardware" claims are postproning BS. What is gonna happen when that happens? They are gonna cry they need GPT-$CURRENT capabilities locally. Now we have local models that are way better that GPT-2 (careful, this one is way too dangerous for release!) GPT3.5, in some ways better that 4, and can run on reasonably modest hardware.
- naasking 6mo agoQuantization can introduce these issues, and Gemma 4 also had issues because the prompt tokens that Gemma used was new and not well supported yet.
- nxtfari 6mo agoI had issues with Qwen thinking endlessly when I didn’t know I wasn’t using the temp/top_k/min_p/etc settings specified in the readme. I’ve never had an issue with Gemma 4 thinking endlessly but could possibly be the same.
- deleted 6mo ago[deleted]
- RITESH1985 6mo ago[flagged]
- tech_curator 6mo ago[dead]
- dzonga 6mo agoif any Alibaba (Qwen) folks are here - website is not working on safari
- smcl 6mo agofuck off: https://news.ycombinator.com/item?id=47796830 https://news.ycombinator.com/item?id=47796830
- poglet 6mo agoCan this run on a PC with 16GB graphics card or a 24GB Macbook Pro? I'm not familiar with how Mixture-of-Experts models differ from standard models.
- bustah 6mo ago[flagged]
- npodbielski 6mo agoI am not sure. I tested it locally on my Desktop Framework and it so far it seem to giving me worse answers then Qwen 3.5. Maybe it is because I am chatting with models in my language instead of enlish or maybe it is optimised for coding instead. I asked it to give me instruction on how to create SSH key and it tried to do it instead of just answering. https://internetexception.com/2026/04/16/testing-qwen-3-6/ https://internetexception.com/2026/04/16/testing-qwen-3-6/
- thesuperevil 6mo agoYeah damn but they are heavy lol
- hemangjoshi37a 6mo ago[dead]
- logicallee 6mo agoWhat kind of hardware does this require to run locally, and how many tokens/seconds does it produce?
- altruios 6mo agoI have moved through the local models at this size. This one is by far the most capable. I've tried various versions of gemma4.26b, various versions of qwen3.5-27/35b (qwopus's galor),nemotron,phi,glm4.7. This one is noticeably better as an agent. It's really good at breaking down tasks into small actionable steps, and - where there is ambiguity - asks for clarification. It's reasoning seems more solid than gemma4, tool use, multi-messaging/longer chain thinking. I am excited to see what other versions of this model people train!
- onlyrealcuzzo 6mo agoHow does it compare to CC Opus Max?
- altruios 6mo agoI try not to use publicly hosted models, and I avoid the SOTA data harvesting machine... so I can not compare. I can compare only to local models. And this one feels like a decently significant leap compared to 3.5 or gemma4. I see there is now a distilled reasoning model version on hugging-face. I may look into that, but I have not seen a need to reach for that change yet either...
- maryjeiel 6mo ago[flagged]
- gck1 6mo agoI have a Macbook M3 Max with 128GB of RAM. How close to Opus 4.6 can I get with this? Realistic, real-world usage. And I mean not sitting there for minutes waiting the model to finish saying hello, or being able to use it for anything more than a pelican riding a bicycle. I'm asking because I'm always seeing excited replies, then I get excited, then I spend minutes to hours setting up the model and then, after first use I forget it exists for one reason or another. Can I get any realistic use out of this?
- stavros 6mo agoYou'd be the best person in this thread to answer this question.
- qazplm17 6mo agoIt won’t be a fair comparison against opus-4.6 but it will run quite well on your machine. I’ve tested qwen3.5 27B, Gemma4, minimax2.5 and Glm4.7 before on my m3 ultra. And i’d say this is the first model that I’m able to use for full agentic sessions. here is a pi session i just did and it worked quite well surprisingly: https://pi.dev/session/#c3d003becb1bfcc7ffbca04e89e1adf8 https://pi.dev/session/#c3d003becb1bfcc7ffbca04e89e1adf8
- gck1 6mo agoThank you! That actually looks quite impressive. What seems very promising is that thinking blocks look coherent for the lack of a better word, and not that far away from thinking blocks (or rather, summaries) that I see from Claude models. I think this could actually work for targeted worker agents that get explicit, detailed task instructions from better models. I'll be trying this tomorrow in my workflow.
- qazplm17 6mo agoJust tried to use qwen3.6-35b-a3b-bf16 + omlx running a pi session to use my HN cli to do a sentiment analysis on this story and opus4.7 story. I’m getting ~40tk/s on a M3 Ultra Mac Studio and the tool use consistency has been held up well. Even when passing 100k tokens, the session was still going strong. Here is the full sentiment analysis report it produced: https://gist.github.com/duh17/2db5351da026cec4bd4f46e169e75e81 https://gist.github.com/duh17/2db5351da026cec4bd4f46e169e75e... Here is the full session: https://pi.dev/session/#c3d003becb1bfcc7ffbca04e89e1adf8 https://pi.dev/session/#c3d003becb1bfcc7ffbca04e89e1adf8 This is by far my smoothest agentic session using a local model of any size. The output quality and speed has really struct the right balance. Very impressive release
- nexustoken 6mo ago[dead]