4 ms·
Insane scores for a model of this size. But it does seem to be a rather insane over-thinker with the biggest token use of any model, which combined with 256k co
by vblanco 1mo ago
Insane scores for a model of this size. But it does seem to be a rather insane over-thinker with the biggest token use of any model, which combined with 256k context size it means it wont do much before filling it context
- spwa4 1mo agoWe don't actually know how much thinking GPT and Opus do, the labs won't show us anymore. And they certainly take their time before starting to answer.
- data-ottawa 1mo agoIt’s definitely a heavy thinker, like most Qwens. I see strings like “write, now.” In the thinning traces then it goes on to think for a lot longer, so it’s kind of weird. I haven’t figured out hope to use this effectively yet on my strix halo.
- cyanydeez 1mo agoWith llamacpp --reasoning-budget truncates thinking and makes it a usefulagain. Its a per client request setting also. Using qwen 3.8 27b, in open code, its unnoticeable. So more a issue for server/harness design than model. Also, there might be a real bug in the models parameters, like this for 3.8-27b: https://huggingface.co/grimoni/Qwen3.8-27B-SSMFIX-UD-Q4_K_XL-GGUF https://huggingface.co/grimoni/Qwen3.8-27B-SSMFIX-UD-Q4_K_XL...