4 ms·
Could you expand more on what you do with qwen3.6? Because I couldn't get the denser 27B version to do trivial "take this pattern, repeat it over a single file
by 3836293648 3mo ago
Could you expand more on what you do with qwen3.6? Because I couldn't get the denser 27B version to do trivial "take this pattern, repeat it over a single file with minimal thought, just slightly beyond what I can do with sed" reliably.
- kgeist 3mo agoHow was qwen3.6 launched? The thing is, everyone has their own variant of "qwen3.6 27b" depending on the launch parameters, ranging from "SOTA in its class" to "completely broken"
- 3836293648 3mo agoYeah, I really should know. But I'm using whatever ollama gave me by default, which probably isn't optimal.
- kgeist 3mo agoOllama uses 4 bit quants and a very short context window by default. It can easily break on anything more complex than a simple chat.
- 3836293648 3mo agoOh, the context window I overrode immediately, the default one can't even reach my prompt in opencode, it's pushed out by all the pre- and post-amble. But I should probably try something less compressed next time I bother.
- anon373839 3mo agoCertainly. First of all, I am using OpenCode as the harness. (I have heard there are better harnesses such as little-coder for small open-weights models, but I haven't tried them yet.) Looking over some of my recent sessions, here are some examples: - Asking Qwen to review project docs (requirements, user stories, etc) so that "we" can evaluate an iterate on an API design. Then back-and-forth chat about possible design directions. Then I ask for a rough-sketch plan of the one I'm interested in. I provide some tweaks to the plan and request a final plan in full detail. I switch to build mode and say go; everything is written to spec. - Asking Qwen to write a suite of tests covering X, Y, Z issues with permutations A, B, C per issue. - Asking Qwen to edit the shape of a CNN to insert auxiliary branches for intermediate supervision, and to extract out part of the network as a modular component with parameterized architecture. I have less experience with the dense 27B because it's too slow to use on Apple Silicon. But regardless of which model you try, I would recommend trying a full-fat cloud hosted version of it first, so that you can get a sense of what it's capable of when the inference stack is correctly configured. LLMs are very sensitive to quantization formats, discrepancies in chat templates, etc. That kind of stuff is make-or-break.