3 ms·
It’s really interesting timing, Qwen over thinking is what kills it for me. I’m just glad we have more options in this size class now.
by wronglebowski 2mo ago
It’s really interesting timing, Qwen over thinking is what kills it for me. I’m just glad we have more options in this size class now.
- dannyw 2mo agoQwen thinking is really good in Mandarin; and probably natively trained the most there. Try a system prompt requiring it to think in Mandarin, while still delivering the response in the user’s language.
- kadoban 2mo agoIs the quality of the thinking better or it's just shorter since Mandarin is more compact?
- yiyu_earth 2mo agoThis is most likely because the vast majority of the information the model absorbed during training was in Chinese. As a native Mandarin speaker, I frequently need to convert the prompt into English and output it in English in order to avoid that the model falls back into Chinese reasoning logic. PS: Switching the thinking process from Chinese to English can also significantly circumvent certain self-censorship mechanisms built into the model.
- ComputerGuru 2mo agoJust to play devil’s advocate: you can’t compare Qwen to a (proprietary/closed source) hosted model and deduce that Qwen is overthinking, as Qwen gives you the full reasoning/thinking trace while all the proprietary models now give you only a summary “to prevent distillation”, making it hard to properly compare apples to apples here.
- naasking 2mo agoPeople say Qwen overthinks because they analyzed the thinking traces, and Qwen finds the answer relatively quickly but then second guesses itself multiple times for another 20,000+ tokens. Regardless of what other models do, that's clearly overthinking.
- seanmcdirmid 2mo agoYou can compare Qwen with thinking to Qwen with no thinking though. I find my results are better without thinking because of overthinking.
- Aurornis 2mo agoYou can tell how long the cloud models spend thinking based on the delay. The Qwen models have a habit of going into thought loops where they go in circles for a while.
- timmmmmmay 2mo agoNo, but you can compare it to the similarly-sized Gemma4 model and see the difference, it's not subtle
- jermaustin1 2mo agoI've been using Qwen3.6 35B A3B, and with reasoning turned on, I'd say 2/3 (give or take) of the tokens for a response are thinking tokens. Which at 70+ tps locally, that isn't that awful. I run an 80k context across 4-10 "agents" for my solo TTRPG, where Qwen is the GM, each NPC at a location, the director, and the narrator. Each turn is about 45-60 seconds to generate all of the various responses. The GM and director have reasoning on, and the NPCs/Location/Narrator do not. It's a fairly good "engine" for that. I'm not sure how a denser Qwen would do here regarding speed.
- jakswa 2mo agoI like the tabletop RPG use case, and wanted to say: If your hardware likes it you should check out Gemma 4 for creative DMing use case. I found it to be much better at holding the plotlines and being creative on gaming turns. My experimental case was an audio-only Zork and Gemma 12B and even E4B were pretty good!
- toyg 2mo agoIs there some sort of dedicated tool for this type of setup, or did you hand-craft it ?
- seanmcdirmid 2mo agoNot parent, but I use Goose for my non-handcrafted Qwen use cases, I’m also working on handcrafting as well. Goose was the only harness that didnt bloat context too much with system prompts (like openclaw) and I could get reasonable web search working with Qwen.
- jermaustin1 2mo agoSomewhat hand rolled, somewhat claude coded. Back in 2023 I started my own C# LLM library for doing tool calls and structured output, and over the years it has morphed bigger and bigger, and that is the backbone of almost all of my LLM-based projects. I've never released it, but its easy to understand, and simple to add your own tools: [AIDescription("Get current weather for a location")] static string GetWeather( [AIDescription("The city name")] string city, [AIDescription("The country name")] string country, [AIDescription("Temperature unit", ["C", "F"])] string unit = "C") { // make some API call to a weather API and return a string to the LLM return $"The weather in {city}, {country} is 22°{unit} and sunny"; } var chat = client.StartConversation("You are a helpful assistant with access to weather data."); var response = await chat.SendAsync<string>("What's the weather in London?", GetWeather); I'm sure plenty of better libraries exist for this now, but in 2023, I don't think any existed in the dotnet ecosystem. I've never released it though, because I've never "finished" it.
- seanmcdirmid 2mo agoDisable thinking? I think many harnesses disable thinking on Qwen anyways because it interferes with tool calling.
- cyanydeez 2mo agoLlamscpp provides reasoning budget and message. You can use the message to redirect it. Once you get the agent and message consistent,itll keep moving.
- ElectricalUnion 2mo agoYou can use any message you want, but the model was tested to react reasonably well to the specific token sequence of "\nConsidering the limited time by the user, I have to give the solution based on the thinking directly now.\n</think>.\n\n" (from a Alibaba paper, struggling to find it now) Edit: arXiv:2505.09388 Qwen3 Technical Report
- cyanydeez 2mo agoSince i have tools to prune context and run subagents, i just tell it to do either since both require summarization which is usually what it needs to avoid the long if...then chains