11 ms·
Qwen3.6-Plus: Towards real world agents
- maxothex 6mo ago[dead]
- srmatto 6mo agoThe benchmarks provided are for Opus-4.5, not for the latest Opus-4.6 and Qwen is still lagging in a lot of them.
- thegeomaster 6mo agoAnd it seems they've decided to go closed-source for their largest, best models.
- FuckButtons 6mo ago3.5-plus was also only available via api. I don’t know what the long term business model for open weights is, I hope there is one, but it seems foolish to assume that companies will be willing to spend millions of dollars of compute on an asset worth zero in perpetuity.
- deleted 6mo ago[deleted]
- vidarh 6mo agoThe business case is to salt the earth for new competitors, coupled with marketing.
- kgeist 6mo agoThey've always had closed-source variants: - Qwen3.5-Plus - Qwen3-Max - Qwen2.5-Max etc. Nothing really changed so far.
- coldtea 6mo agoThey always did that. Did they say anywhere they'd open all their models? They still have a business.
- Aurornis 6mo agoThere is no reason to benchmark against Opus 4.5 when Opus 4.6 has been out so long, other than to be misleading.
- coldtea 6mo agoI can see reasons, among others that 4.5 was the one established as they were preparing this version. "So long" is merely 2 months ago, and Qwen 3.5 was barely released less than 2 months ago. They were likely already working on finalizing 3.6 before 3.5 official launch, and as 4.6 came out. In any case, aside Claude fanboyism, having other plays inch closer to similar performance is always useful. Even if they are "6 months behind" as the pace slows down, this guarantees that there's no huge moat and they'll eventually either get to where the SOTA is, or the difference wont be that big. I'd rather put fewer eggs in 2-3 big player baskets.
- jgbuddy 6mo agoWorth noting that this model, unlike almost all qwen models, is not open-weight, nor is the parameter count exposed. Also odd that it is compared against opus 4.5 even though 4.6 was released like 2 months ago.
- pferdone 6mo agoThey said in the last paragraph[0]: "[...] In the coming days, we will also open-source smaller-scale variants, reaffirming our commitment to accessibility and community-driven innovation. [...]" [0] https://qwen.ai/blog?id=qwen3.6#summary--future-work https://qwen.ai/blog?id=qwen3.6#summary--future-work
- deaux 6mo ago> we will also open-source smaller-scale variants In other words, like GP said, this Qwen3.6-Plus model is not open-weight unlike the other Qwen models.
- pferdone 6mo ago> unlike almost all qwen models Almost all means there have been ones before that were not open. So, no contradiction there.
- kennywinker 6mo ago> unlike the other Qwen models Please send the download link for qwen 3.5-plus. Also, who cares? If you have the hardware to run a ~400b model i don’t think you count as a home user anymore.
- dgb23 6mo agoIn a practical sense, I'm primarily interested in small to medium sized models being open. I think that might be common sentiment. However, my hope is that there will be at least somewhat competitive big and open models as well, from an ethical/ideological perspective. These things were trained on data that was provided by people without their consent, so they should at least be be publicly accessible or even public domain.
- Art9681 6mo agoHow convenient of them to compare themselves to the last generation Opus and GPT models to make their model look better than it really is.
- MarsIronPI 6mo agoIt's not open weights so I'm not interested.
- karimf 6mo ago> In the coming days, we will also open-source smaller-scale variants, reaffirming our commitment to accessibility and community-driven innovation.
- Aurornis 6mo agoThis is their hosted-only model, not an open weight model like they’ve become known for. They got a lot of good publicity for their open weight model releases, which was the goal. The hard part is pivoting from an open weight provider to being considered as a competitor to Claude and ChatGPT. Initial reactions are mostly anger from everyone who didn’t realize that the play along was to give away the smaller models as advertising, not because they were feeling generous. Comparing to Opus 4.5 instead of the current 4.6 and other last-gen models is clearly an attempt to deceive, which isn’t winning them any points either. I think there is a moderately large market for models like this that aren’t quite SOTA level but can be served up much cheaper. I don’t know how successful they’ll be in the race to the bottom in this market niche, though. Most users of cheap API tokens are not loyal to any brand and will change providers overnight each time someone releases a slightly better model.
- cubefox 6mo ago> I think there is a moderately large market for models like this that aren’t quite SOTA level but can be served up much cheaper. There isn't, pretty much everyone wants the best of the best.
- scoopdewoop 6mo agoThat isn't true. In a Codex or Claude Code instance, sure... but those are not the main users of APIs. If you are using LLMs in a service for customers, costs matter.
- Aurornis 6mo agoThe market for API tokens is bigger than people like you and I (who also want the best) using then for code. There are a lot of data science problems that benefit from running the dataset through an LLM, which becomes bottlenecked on per-token costs. For these you take a sample subset and run it against multiple providers and then do a cost versus accuracy tradeoff. The market for API tokens is not just people using OpenCode and similar tools.
- sidrag22 6mo agomaybe there isnt, but as understanding grows people will understand that having an orchestration agent delegate simple work to lesser agents is significant not only for cost savings, but also for preserving context window space.
- daft_pink 6mo agoNot really interested in using models hosted on alibaba cloud. Like Qwen local for it’s privacy, but I trust the privacy of Google/OpenAI/Anthropic more than alibaba.
- rvz 6mo ago> Like Qwen local for it’s privacy, but I trust the privacy of Google/OpenAI/Anthropic more than alibaba. None should be trusted, unless you are running them locally.
- the_pwner224 6mo agoI had the exact opposite reaction. I stopped using OpenAI/Google a while ago due to privacy and moved to local Qwen, now I'm considering using Alibaba cloud. You know Google and OpenAI are going to share everything with the US government and Western ad networks. But with Alibaba, who cares if the CCP & Chinese ad networks have a comprehensive profile on me? From a pragmatic perspective it's much better for (outcomes related to) privacy.
- zobzu 6mo agoso if China has the data good, us has the data bad, got it lol. us actually has laws around this and they arent sharing very much with thr us gov today. china shares 100% as required by law. and neither care much about "how long do i cook eggs for", but they do care about code generation a lot.
- thereitgoes456 6mo ago> so if China has the data good, us has the data bad It's not that, it's about relative risk to your own life. Asking questions about "DEI" for example is much more likely to have adverse effects on your life if you ask Grok or an OpenAI chatbot, though still not that likely.
- wongarsu 6mo agoFrom an espionage perspective your own government is the safest. But from a civil rights perspective your own government is your most immediate threat. China isn't going to arrest me for my opinions on Netanyahu, my own government could And the US government has repeatedly shown that it is very interested in collecting all the data available, just like China. In China this is simply done in the open while the US has a veneer of protection for citizens. But where the data collection is forbidden by law they either ignore the law or ask another five eyes member to do the spying and share the results. Both are well documented
- woeirua 6mo agoJust more evidence that the B tier models are six months behind. Ultimately that’s good. Opus 4.6 level intelligence will be cheap later this year!
- eis 6mo agoQuite strong results in the benchmarks but why Gemini 3 Pro instead of 3.1? Why only for a few of the benchmarks? Why is OpenAI not there in the coding benchmarks? Why Opus 4.5 and not 4.6? Just jumps out into my eye as a bit strange. As always, we'll have to try and see how it performs in the real world but the open weight models of Qwen were pretty decent for some tasks so still excited to see what this brings.
- Caum 6mo ago[flagged]
- rogerrogerr 6mo agoIn my experience, whenever an LLM says “wait, but actually” or some variant thereof is when you need to step in before it goes totally off the rails.
- esafak 6mo agoDoes anyone have experience with Alibaba's coding plan? Not that I'm very tempted at $50/month...
- usagisushi 6mo agoA bit off-topic but I’m on the legacy Lite plan (now discontinued), and it’s more than enough for hobby projects. The main draw is the generous request-based quota (18k requests/month) rather than a token-based one. This means a 100k token request counts the same as a 100-token one. I’ve made about 8000 requests in the last two weeks, averaging around 80k tokens per request. It feels like they’re subsidizing this just to gather data on agentic workflows. On the downside, the speed is mediocre (15–30 tg/s for GLM-5), and I’ve seen the model glitch or produce broken output about 10 times out of those 8k requests.
- linolevan 6mo agoI’m surprised that people are surprised. Qwen has been hosting private plus and max variants for a while now.
- fdsjgfklsfd 6mo ago> In particular, Qwen3.5-Plus is the hosted version corresponding to Qwen3.5-397B-A17B with more production features, e.g., 1M context length by default, official built-in tools, and adaptive tool use. For more information, please refer to the User Guide. https://huggingface.co/Qwen/Qwen3.5-397B-A17B https://huggingface.co/Qwen/Qwen3.5-397B-A17B So 3.5-plus has been released as open weights.
- giancarlostoro 6mo agoI hope their open source variants are just as good, having a 1 million token window for a fully offline model would be VERY interesting.
- sosodev 6mo agoI don't know how well it performs, but you can extend Qwen3.5 to 1 million token context using YaRN. Also, Nemotron 3 Super was recently released and scales up to 1 million token context natively.
- throwaw12 6mo agoI would love to hear from people using both (Claude Code OR Codex) AND (Qwen) and their experience with Qwen models, are they on par, or how far are they?
- scottcha 6mo agoI switch between Claude Code (Opus/Sonnet) and Qwen (OpenCode, OpenClaw) multiple times throughout the day and Qwen 3.5 is really nice. I do also use KimiK2.5 and GLM5 pretty often too and I'm starting to get a sense that the agent tool is becoming a little more important than the model with these level of models. As long as tool calling and prompt quality is all configured correctly by the provider.
- danelliot 6mo ago[dead]
- edg5000 6mo agoWhat are the reasons for switching? Personally I got into the habit of doing a bit of a round robin with Codex/Claude (CLI) and then DeepSeek and Qwen web chat. And Claude in web chat. I like to switch just to learn the differences, otherwise I'd never know what the other models can do. But I still feel attached to Opus, but this can be fammillarity. If I only had Qwen maybe it would be effectively identical at the end of the day. Hard to say.
- scottcha 6mo agoMine are pretty unique since we optimize the energy for and run an inference service api so forces me to dogfood alot of different options.
- avib99 6mo ago[dead]
- fdsjgfklsfd 6mo agoQwen3.5-plus is quite good, Qwen3.6-plus is not.
- furyofantares 6mo agoI'll diverge from some of these comments, I don't find it misleading to compare to Opus 4.5. I can remember how good Opus 4.5 was. If I'm considering using this, it's most informative to me to compare to the model it's closest to that I have familiarity with. I'm obviously not switching to this if I want the best model. I'm switching if I'm hopeful that the smaller versions are close to it, or if I want to have more options for providers, or for any other reasons unrelated to getting the highest quality responses possible.
- bensyverson 6mo agoExactly this. If you can get something close to Opus 4.5 for free, that's noteworthy. I may not use it for the most critical pieces of my app, but not everything I do is galaxy-brain coding.
- cmrdporcupine 6mo agoYes, honestly, Opus 4.6 and GPT 5.4 were mostly not really noticeable improvements over 4.5 and 5.3 respectively. If we were stuck at 4.5 levels but at 1/10th of the price, I'll take it.
- furyofantares 6mo agoI find 4.6 pretty noticeable upgrade, but it might be the 1M context. I'm interested in how the 1M context works out with Qwen.
- nwienert 6mo agoI found it worse, in a very clear way.
- Alifatisk 6mo agoFrom Qwen-3-max thinking, I remember the inference becoming veeery slow as you pushed towards 1M context, already at 300k tokens you would notice the degradation. But of course, I was using Qwen Chat, so could be a resource allocation thing.
- Alifatisk 6mo agoI understand peoples reactions of Qwen team comparing against Opus 4.5 instead of 4.6. And them comparing against Gemini Pro 3.0 instead of 3.1. But calling it misleading is a bit of stretch in my eyes, people here are acting like we immediately forgot how previous generations performed just because a new version is released. This field is going in a incredible pace, the providers release a new model every quarter or so. The amount of criticism is a bit overblown in my opinion. The benchmarks still look very good to me. I’ve used GLM-5 (latest is GLM-5.1) and Kimi K2.5, they are decent and gets the job done, so seeing how this model of Qwen performs compared to it is kinda impressive. Also, why are so many pointing out the fact that this model is not open-weight as if this is their first time doing so. Qwen-3.5-plus, Qwen-3-Max is also closed source. This is not something new. I think Qwen trying to catch up to the SOTA models is still healthy for us, the consumers. Sure, its sad news that this version is closed-weight, but I won’t downplay their progress.
- nickvec 6mo agoI think it’s more the principle of deception that upsets people. Imagine if Apple released a new iPhone and publicly compared its specs to some previous gen Android. It’s not in good faith.
- Alifatisk 6mo agoWhy are we so quick to call it deception? Their figure is quite clear. They aren't fiddling with the graph or hiding the labels, they are clearly stating which models it compares against. But I agree on the sentiment that the standard practice should be to bench against the latest SOTA models.
- deleted 6mo ago[deleted]
- patates 6mo agoEven if openly stated, why would they be comparing to a previous generation if not for deception? Laziness? Lack of time? It's not like the latest generation of the SOTA models were released yesterday.
- techpulselab 6mo ago[flagged]
- zkmon 6mo agoIt is no longer available on OpenRouter. They say "going away on 3-March", but it's already gone!
- simonw 6mo agoPretty solid Pelican: https://gist.github.com/simonw/ca081b679734bc0e5997a43d29fad879#file-pelican-svg https://gist.github.com/simonw/ca081b679734bc0e5997a43d29fad... I used the https://modelstudio.alibabacloud.com/ https://modelstudio.alibabacloud.com/ API to generate that one, which required signing up for an account and attaching PayPal billing - but it looks like OpenRouter are offering it for free right now so I could have used that: https://openrouter.ai/qwen/qwen3.6-plus:free https://openrouter.ai/qwen/qwen3.6-plus:free
- bredren 6mo agoPelican is drafting rear peloton
- manc_lad 6mo agothey're going to start training a pelican riding a bike specifically on these models soon. it's the key global benchmark!
- teruakohatu 6mo agoThey will have a long time ago. By now Simon's meme will be well represented in training sets.
- wg0 6mo agoIt hallucinates a lot more then Sonnet or even MiniMax M2.5. Especially in tool calls, it would end up duplicating the content in code files and then realising later and getting stuck in a loop.
- noelsusman 6mo agoMy initial experiments are not encouraging. I have a basic planning prompt that includes instructions not to edit any files or implement anything. Qwen-3.6-Plus will consistently ignore that completely and proceed with implementation. I expect that kind of behavior from small models I run locally, not a hosted closed model claiming to compete with the frontier models.
- justinclift 6mo ago> It hallucinates a lot more then Sonnet or even MiniMax M2.5. Ugh, that's not good. I evaluated Kimi K2 a while back for some text understanding -> summarisation tasks, and of the 100 tasks it hallucinated about 30% of the output. :( :( :(
- dryarzeg 6mo ago> I evaluated Kimi K2 a while back I guess that it was Kimi K2-Instruct, the first model (or it's fine-tune) in the lineup of Kimi-K2 models. And I remember trying it just for the sake of curiosity, and... except for the almost total absence of the sycophancy and "sugar syrup" in it's outputs, it was not very good at the time. Right now though, if you're still interested in this model family, you could look at Kimi-K2.5 which is way better. That said, it's still not perfect, and to be honest, looking where things are going with LLMs right now I prefer the use of my own brain (local private inference with power consumption of ~20-25W, having a capability for continuous learning and performing real-world tasks) to the use of any "AI" model (including proprietary models such as Claude 4.6 Opus, Gemini 3.1 Pro and others). : )
- wolvoleo 6mo agoNice, I hope there will also come a small open version of it.
- shubhamgarg86 6mo ago[flagged]
- kanehorikawa 6mo ago[dead]
- johnwhitman 6mo ago[flagged]
- gburgett 6mo agoLooking forward to when this gets on Bedrock. I built an app with a niche AI agent and to this point only Sonnet is really good enough for our use case, but its expensive!
- dzonga 6mo agoQwen free plan is still good. you get a generous token limit.
- davesque 6mo agoI wish these AI vendors would quit publishing comparisons with the previous generation of their competitors's models. It's just such a glaringly bad look and no one is fooled by it, even if their achievements deserve praise in their own right. The Qwen models are great and don't deserve the reputational hit that comes from dodgy marketing tactics.
- XCSme 6mo ago3.6 Plus seems to be simply a refined/more consistent 3.5 Plus: https://aibenchy.com/compare/qwen-qwen3-5-plus-02-15-medium/qwen-qwen3-6-plus-medium/ https://aibenchy.com/compare/qwen-qwen3-5-plus-02-15-medium/...
- adinhitlore 6mo agoi've been fan of qwen for quite some time, awsome!
- try-working 6mo agoFor anyone that believes Chinese labs will stop open sourcing their models, let me tell you why that won't happen. First, try signing up for Z.ai's coding plan. I know how to but I bet you won't be able to. The absolute disaster that is Z.ai's internet presence shows that these small labs have no ability to market themselves and drive direct sales. For marketing, they lack capabilities, and releasing open models is the only way for them to remain in the conversation. For sales, they rely on distribution via OpenRouter, OpenCode etc. Interest with their users is driven by open model performance. Open sourcing for Chinese labs is not some large national scheme. It is their only way to commercialization.
- rogerrogerr 6mo agoWell, can’t they just direct their model to do some marketing for them? Only partially tongue in cheek - if it’s not good at marketing itself, that seems like a red flag for capabilities?
- polski-g 6mo agoChina can't get good chips. But I don't understand why they can't license their closed source models to US inference providers so we can get more than 80% reliability on their models on OpenRouter.
- try-working 6mo agoI think they already are licensing their biggest models to third party inference providers.
- sumedh 6mo agoAgree, it’s surprising that these companies don’t thoroughly test their own workflows to ensure a smooth and seamless user experience.
- Alifatisk 6mo ago> First, try signing up for Z.ai's coding plan. I know how to but I bet you won't be able to. What's the issue with signup up for Z.ais coding plan?
- volume_tech 6mo ago[dead]
- zwaps 6mo agoThey claim SOTA but are beaten by last gen Opus om every metric? This one seems weird
- Sim-In-Silico 6mo ago[dead]
- kristopolous 6mo agoI've gone through about 500M tokens on this model already. They've got some free inferencing options (such as on openrouter) ... $0 is hard to beat and it's creating not-crap.
- edg5000 6mo agoHow can it be free? What do you mean? EDIT: Ah, I see. Some kind of promotion. Pretty cool.
- kristopolous 6mo agoAlso the Qwen cli, their vibe coding agent, allows a thousand free requests a day. In practice it's about ~300M tokens for me It's very generous I know I could certainly pay anthropic $1500 a day for my use and they'd be delighted... I'd rather pay $0 ... Just a personal preference
- edg5000 6mo agoHahaha
- nl 6mo ago23/25 on my agentic benchmark for the free version on OpenRouter. That's a great score - only 4 models have ever scored higher. But there are open models that also score 23/25 including Qwen 3.5 27B.
- edg5000 6mo agoHas anybody done serious agentic work (e.g. using a CLI harness or simmilar) with 3.5 Plus/3.0 Max and such? How does it compare against Opus with Claude Code? I've used the chat quite a bit and I can't say at this point.
- fdsjgfklsfd 6mo agoQwen3.5-plus has been my go-to model for non-autonomous chatbot that runs arbitrary code and shell commands on my local machine for one-off tasks (unlike Claude Code). It has no problem calling tools to run code (that contains tools), and pretty smart, and cheap (though lack of token caching is making it much more expensive than it should be). I tried Qwen3.6-plus and it was not as good.
- edg5000 6mo agoDo they have an API where you can control the chat template or at least just put everything in the system prompt? This way you can control everything including the tool calling syntax. Even if you use the trained tool syntax, it allows you to control the tool system prompt which you may want to tweak. With DeepSeek this is all possible. An undocumented feature, great for harness builders. Anybody got info on Qwen regarding this?
- Aerroon 6mo agoCan't you do that on OpenRouter? You can set a system prompt there. Is that insufficient for what you had in mind?
- breisa 6mo agoThey are talking about the chat template and not the system prompt. With current gen models, the system prompt is only part of a larger pretext that is passed to the model at the start of the "chat". The models are trained on a specific chat template with things like tool lists, reasoning budget, special feature flags and the "system prompt" formatted in a certain template.
- Vivolab 6mo ago[dead]
- throwaway911282 6mo agoignoring gpt 5.4! I feel bad for people who have not even tried it. for the same 20$ I pay to openai and anthropic, I get significantly more from openai
- EdoardoIaga 6mo ago[flagged]
- pratyushsood 6mo ago[dead]
- 0xqlive 6mo ago[dead]
- geenkeuse 6mo ago[dead]
- mtrifonov 6mo agoMost agent work focuses on task completion. Browse the web, fill out the form, and/or write the code. The harder problem is social agency, where the AI has to decide whether to participate at all. We built a cheap model gate that reads the conversational dynamics of a group chat before the expensive model runs. Wonder how Qwen3.6 performs in these nuances cases.
- bezlant 6mo ago[dead]