2 ms·
I will TLDR you on our thought process, research, training and benchmarks. 1. When running our quantized Qwen 3.8 27B instances we were very annoyed by random
by kisjovan 19d ago
I will TLDR you on our thought process, research, training and benchmarks.
1. When running our quantized Qwen 3.8 27B instances we were very annoyed by random reasoning loops (in the paper bellow refered to as "overthinking errors". These random loops were persistent throughout medium and low reasoning settings.
2. We found a paper by Meta that's supposed to target this phenomenon in PTQ, but when used straight out of the box got mixed results. https://arxiv.org/abs/2606.00206 https://arxiv.org/abs/2606.00206
3. We figured to try if it's a matter of the targeting the right keywords and tuning the parameters, so we used our 8xH100 box and and generated a large amount of different (ofc out of distribution) domain (coding, language, vision, agentic) traces.
4. We then grouped the ones with overthinking and found "common denominator" tokens between them and targeted the most prominent ones.
5. We then built an inference-time penalizer of those tokens as seen in the paper with the hopes of simply generating traces and doing cross-entropy SFT over them.
6. Did not work at all, but the penalizer seemed to work much better than the tokens provided in the paper and not only for lower precision models but for bf16 as well. Hence we kept experimenting with it. We built a loss function using the tokens we identified and ran LoRa SFT over the traces prev generated and reasoning seemed to be falling off significantly but the accuracy seemed to follow. The reasoning reduction seemed to be generalizing.
7. After a significant amount of tinkering (literally since the day of Qwen 3.8 27B release) we were satisfied with the reasoning reduction. After that we searched for ways of restoring the accuracy. We experimented with several methods, including RL(GSPO), On-Policy Distillation and using the ThinkingCap 3.6 27B adapter chunks until we were satisfied with our accuracy loss. We managed to restore it to <1% loss on almost all of our OOD in house tests
8. We then performed intensive intensive benchmarks, across several reasoning efforts, precision variants etc. We ran into a few problems, one of which is that to get a reliable score we needed to run each benchmark 10x (5x on base + 5x with our adapter, this being the standard procedure on the Qwen 3.6 27B model card on Terminal Bench which we followed). After running it, the performance converged to 40-60% token reduction with <1% accuracy loss across GPQA, MMLU, Terminal Bench 2.1, LiveCodeBench v6, ERQA, C-Eval, IFBench, HMMT25, with an exception being AIME26 with an accuracy loss of 4.6%, which we later linked to a bug during training with a specific token relevant for math-related reasoning being penalized and are planning to fix it in an updated release.
- pu_pe 19d agoCongrats, seems like a promising concept. I'll give it a go in my local setup. My prior is thinking this will definitely work for speed but also definitely compromise accuracy (I've seen this happen so many times), though benchmark results are encouraging of course. Nice explanation too. Now that everything reads like a sad blur, it felt refreshing to read your writeup.
- kisjovan 18d agoThank man appreciate it! Did you get the chance to try it?
- alpha_trion 19d agoNice. Going to dl and give it a whirl.
- kisjovan 18d agoDid you get to try it?
- billziss 19d agoI tried an oMLX quant of your model (suzu89/Swift-Qwen3.8-27b-oQ8-mtp -- not mine) and liked it. It certainly seems to cut down on thinking compared to stock Qwen3.8 27B in my (limited) testing. A couple of questions: - Have you tried the peculiar-ragdoll/Qwen-Sharp-Chat-Templates with it? They replace the default chat_template.jinja with one that encourages less thinking. - When are you releasing Swift-Qwen3.8-Flash-Next? :)
- kisjovan 18d agoThank you so much for trying it! I personally haven't tried it - the community has noted that it does indeed work though. On the Swift Qwen3.8-Flash-Next, we're running the benchmarks right now and will get it out end of this or start of next week! You'll for sure find it on r/LocalLlama