20 ms·
I tried an oMLX quant of your model (suzu89/Swift-Qwen3.8-27b-oQ8-mtp -- not mine) and liked it. It certainly seems to cut down on thinking compared to stock Qw
by billziss 11d ago
I tried an oMLX quant of your model (suzu89/Swift-Qwen3.8-27b-oQ8-mtp -- not mine) and liked it. It certainly seems to cut down on thinking compared to stock Qwen3.8 27B in my (limited) testing.
A couple of questions:
- Have you tried the peculiar-ragdoll/Qwen-Sharp-Chat-Templates with it? They replace the default chat_template.jinja with one that encourages less thinking.
- When are you releasing Swift-Qwen3.8-Flash-Next? :)
- kisjovan 10d agoThank you so much for trying it! I personally haven't tried it - the community has noted that it does indeed work though. On the Swift Qwen3.8-Flash-Next, we're running the benchmarks right now and will get it out end of this or start of next week! You'll for sure find it on r/LocalLlama