3 ms·
So are locally-runnable models frozen at Qwen 3.6 now :/
by ernsheong 3mo ago
So are locally-runnable models frozen at Qwen 3.6 now :/
- worldsavior 3mo agoEveryone wanted open models that would challenge Opus and Codex, here, you got it.
- ernsheong 3mo agoWe need better coding models that can run on local hardware, i.e. 128GB VRAM or less
- seanmcdirmid 3mo agoQueen has that already, although they seem to be moving away from local models unfortunately.
- zozbot234 3mo agoYou can run larger models by offloading to SSD (for weights), it's just slow so people don't do it all that much. But you can get back at least some of that performance by using either MTP (at least for dense models; not effective for sparse MoE models unless you're batching them already and have VASTLY more parallel compute than you'd know what to do with) or batching multiple requests in parallel (note, this hurts throughput for your single sessions but running more sessions in parallel still boosts your total amount of inference. This requires careful management of memory requirements for your context/KV cache, and Qwen models tend to be KV-cache heavy). Broadly speaking, this ultimately pushes local inference towards a challenging world where you use SSD offload for weights as a matter of course; then smaller requests (or requests sharing the bulk of their context, e.g. subagent swarms) can be batched together and run quickly in aggregate, but running very large contexts will actually limit you to single-session inference and require swapping out even the KV cache itself to some external scratch SSD, further hurting your performance. Then feel free to add wide use of MTP in a probably futile effort to go back to tolerable tok/s numbers.
- tormeh 3mo agoIs qwen 3.6 27b the best model you can run locally at the moment? Not that I have the VRAM for it, but just curious.
- cmrdporcupine 3mo agoGemma4 models are arguably better. Or at least about the same.
- ch_sm 3mo agoIn my experience, yes. A bit more reliable than gemma for me. I mostly use A3B (35B, mix of experts) though, because it‘s faster, and in the same ballpark intelligence wise as the dense 27B, so it’s the sweetspot for me. I want to try cohere‘s mini code model next, but worried the runtimes aren‘t optimized for that yet.
- mark_l_watson 3mo agoI found qwen3.6:26b slightly better on my 32G mac mini than the same sized gemma until gemma was updated with better tool support 4 or 5 days ago. It is like a ping-pong game: the advantage flips back and forth between providers.
- regularfry 3mo agoWorth knowing that Unsloth have just put out another Gemma 4 release from Google's upstream updates which should improve reliability. Bugs in the chat template affecting tool calling and other issues, apparently. https://www.reddit.com/r/unsloth/s/MpC6Hzs4Wj https://www.reddit.com/r/unsloth/s/MpC6Hzs4Wj
- dofm 3mo agoWow, thanks. I didn't see Unsloth had already done their version; I was just about to go back to the google version to test this change.
- regularfry 3mo ago
- deleted 3mo ago[deleted]
- dofm 3mo agoMaybe, maybe not. Qwen 3.6 27B is literally just three months old. Hard to predict. Maybe it just wasn't worth making a 3.7, and after all, the 27B release was after the Plus release.