2 ms·
What's going on with Qwen3.6 27b? Filtered to Python it comes out at the top of the list, which seems... well, unlikely.
by regularfry 5mo ago
What's going on with Qwen3.6 27b? Filtered to Python it comes out at the top of the list, which seems... well, unlikely.
- 2ndorderthought 5mo agoQwen3.6 27b is a really strong model.
- regularfry 5mo agoYeah but that strong?
- 2ndorderthought 5mo agoYes that strong. Its only lacking in context length, but it's not that small there and it gets caught in circles more often then say a 1t parameter model does. That's why a lot of people have been freaking out about local LLMs since april. There's finally a decent model that runs locally on a GPU or two that can do agentic programming at a reasonable enough tokens per second.
- johndough 5mo ago> it gets caught in circles more often then say a 1t parameter model does. I've found that the Q5+ quants are less loopy than Q4. Still not perfect, but noticeably better. > reasonable enough tokens per second The speed has been amazing. I've been running the recent llama.cpp MTP branch with an uncensored variant of Qwen3.6-35B-A3B on my RTX 3090 over 170 tokens per second and it was able to turn a buffer overflow into a reliable shell exploit in just a few seconds (with reasoning disabled). Still a bit loopy though. Hopefully, the Qwen team will pay more attention to those looping issues. It feels like their models are especially susceptible.
- 2ndorderthought 5mo agoIs that on a single 3090? I need to change my settings it sounds like
- johndough 5mo agoYes, single RTX 3090 with this model https://huggingface.co/llmfan46/Qwen3.6-35B-A3B-uncensored-heretic-Native-MTP-Preserved-GGUF/blob/main/Qwen3.6-35B-A3B-uncensored-heretic-Native-MTP-Preserved-Q4_K_S.gguf https://huggingface.co/llmfan46/Qwen3.6-35B-A3B-uncensored-h... following these https://huggingface.co/havenoammo/Qwen3.6-35B-A3B-MTP-GGUF https://huggingface.co/havenoammo/Qwen3.6-35B-A3B-MTP-GGUF instructions (should add "-j 8" to last cmake command for parallel build) and llama-server with --reasoning off Note that the MTP PR https://github.com/ggml-org/llama.cpp/pull/22673 https://github.com/ggml-org/llama.cpp/pull/22673 is still under development, so things might be broken.
- johndough 5mo agoWhile Qwen3.6 27B and 35B-A3B are very good, I am skeptical about them being that good. I think another factor is at play here. The Qwen3.6 models have memorized some common games. For example, if you ask it to create an index.html with a snake game, it will generate almost the same high quality snake game every time. The relatively low success rate of 25% but high average percentile of almost 100% for one-shot coding in Python suggests that the model is extremely good at few tasks.
- gertlabs 5mo agoThe more filters applied (one-shot coding only, Python only), the more variation you can expect from fewer samples -- that being said, it really is a great model so it's probably not too far above where it would end up with infinite samples.