4 ms·
I used Qwen 3.6 27B extensively (>1B tokens) and DeepSeek V4 Flash (the older one also 2B+ tokens). And I just can't fathom that the new 3.8 beats the new Deep
by K0IN 2mo ago
I used Qwen 3.6 27B extensively (>1B tokens) and DeepSeek V4 Flash (the older one also 2B+ tokens).
And I just can't fathom that the new 3.8 beats the new DeepSeek V4 Flash (which, in my eyes, is one of the best everyday coding models).
What an insane release, and convenient size to use every day/locally.
but i will test this model extensivly.
- JacobAsmuth 2mo agoIt has double the active params.
- drob518 2mo agoI’ve been using v4 Flash 0731 a lot lately and you can’t beat the price performance. That said, it sometimes takes my prompts as more of a suggestion than a directive. I’ve found that introducing a reviewer subagent (even with the same model) helps push it back to what I’ve asked for. But makes every coding session a back and forth: “do X” -> “use a reviewer subagent to analyze whether you really did X as I asked”.
- Saris 2mo agoWhat model do you normally run the subagent on? You mentioned flash as well for that, but I wonder if a more 'strict' model would do a better job at pushing the main back on track.
- drob518 2mo agoFor cost reasons, I’ve been using Flash for the reviewer, too, but I plan on trying to use Pro for that. Thus far, however, Flash has been doing well at reviewing. I’m cheap as I’m paying for all the tokens myself.
- 0xc133 2mo agodeepseek-v4-flash-0731 has been awfully prone to infinite looping output in reasoning for me, and once it hallucinated in the middle of going in circles that I had instructed it to start using Yoda-speak, which… I don’t have any idea where that came from.
- drob518 2mo agoI have had it loop once. I had to abort the turn, clear the context (maybe I just compacted it?) and restarted. That seemed to clear it up. In general, I have not found that V4 Flash 0731 loops a lot or even any more than other models.
- 0xc133 2mo agoHow were you accessing the model? I was using ollama.com’s cloud hosting the first few times I ran into it, and then it happened again the first time I tried Fireworks.ai before I made it even 10% of the way through the context window.
- drob518 2mo agoPi -> OpenRouter -> Deepseek V4 Flash running on a random 3rd party provider (not Deepseek itself)
- f311a 2mo agoHow is the general knowledge of Qwen 3.6? Do you need to explain things outside of algorithms to it? Since the size is so small, I guess you need more explanations to it. General knowledge helps with coding when your don't specify a lot of details and ask for big changes.
- SwellJoe 2mo agoIt researches what it doesn't know, just give it a web search tool. It searches unprompted, if it can. It's impressive.
- algo_trader 2mo ago> I used Qwen 3.6 27B extensively (>1B tokens) and DeepSeek V4 Flash Were your opinions effected by the harness ? DS is an amazing combo. It probably could only happen in China, not in current USA or EU (for different reasons)
- ignoramous 2mo ago> It probably could only happen in China, not in current USA or EU (for different reasons) Per Artificial Analysis benchmarks, Meta's Muse Glimmer 30b (open weight) holds its own (for agentic code workloads) against models 5x to 10x its size, too.
- kube-system 2mo agoand Muse Glimmer uses a lot fewer tokens than Qwen 3.8. I found it more usable on my hardware because I can get an answer quicker.
- skohan 2mo agoI found Glimmer underwhelming in terms of coding - I tried it as a drop-in replacement for 3.6, and the output was noticeably worse. 3.8 has been a significant step up so far from early testing.
- skohan 2mo agoGlimmer benchmarks around Qwen 3.6 27B levels no?