10 ms·
Qwen3.7-Max: The Agent Frontier
- bsenftner 5mo agoAny reports from people using their coding agent(s)?
- rayboy1995 5mo agoI'm running Qwen 3.6 27B Q5 K M GGUF on a Tesla P40 and koboldcpp using pi.dev as the harness, I gotta say I am impressed. Took some setup and configuring but I already have some code it has made commited and pushed. It can be slow on my hardware at >50k tokens, but the fact I bought this one P40 for like $150 back when the LLM trend started I can't complain. (I have a second one too but I couldn't physically fit the card in my server unfortunately.) The setup I had to do was important and I had to compile koboldcpp with a few special params for my hardware, I mostly just had Claude figure it out. I don't remember everything I did now but it was very slow and would often stop mid task, it seems it was mostly a parsing issue. It made the model seem broken/dumb, but once I had all that settled I actually am able to use this how I use Claude Code. Disclaimer, I am pretty explicit with requirements, I imagine this fails more when you leave it to figure out things on its own but for my flow its pretty rad. Currently setting it up as an automated agent now to pull Trello cards, create PRs for them, and move the card to be reviewed. Command I am using to run: python koboldcpp.py \ --port 61514 --quiet --multiuser --gpulayers 999 --contextsize 262144 --quantkv 2 \ --usecublas normal --threads 4 --jinja --jinja_tools --jinja_kwargs '{"enable_thinking":true, "preserve_thinking":false}' \ --skiplauncher --model /data/models/Qwen3.6-27B-Q5_K_M.gguf --smartcache 5
- lostmsu 5mo agoQwen recommends to preserve_thinking: true for agentic/coding workloads.
- rayboy1995 5mo agoThanks!! I had disabled that previously while debugging, I can confirm this is helping accuracy from what I can tell so far. (And speed since the cache is preserved more often!)
- satvikpendem 5mo agoUse the MTP models which 2x token generation speed, for example: https://unsloth.ai/docs/models/qwen3.6#mtp-guide https://unsloth.ai/docs/models/qwen3.6#mtp-guide
- rayboy1995 5mo agoVery interesting I'll have to check this out thank you. This is why I love HN.
- vibe42 5mo agoI'm using the pi-mono coding agent (open source, free) without any extensions and very simple prompts. The 3.6 27B model (BF16, 250k context) uses 67GB VRAM on an RTX PRO 9000. It's very capable on almost any coding task I've thrown at it, and very good for easy-to-medium hard scripts, new code bases. It struggles on some complex tasks in larger code bases, e.g. using to debug and fix bugs in llama.cpp it gets close to working code but often introduces errors. For such tasks its still very useful as a search/explore tool and drafting fixes.
- jdw64 5mo agoQWEN really hits the sweet spot it's cheap, fast, and actually good.
- goyozi 5mo agoThese are very good numbers. I still don’t get why they don’t compare against latest competitor versions in these posts, it’s not like we’re all not going to notice.
- hmokiguess 5mo agothis puzzles me too, I want to know
- htrp 5mo agoI think its part of the expectation setting (with a side of we did our distillation/ eval harness on a specific model). if they say it's 4.7 comparable, it anchors that into your head as the model to evaluate against.
- maelito 5mo agoMarketing.
- Aurornis 5mo agoI think the argument is that trying to suggest that they’re close to N months from SOTA. Realistically I assume they hope readers don’t notice the fine details. The Qwen models are great for open weights but for every past release they haven’t performed as well as the benchmarks in my experience. They’re optimizing for benchmark numbers because they know it works.
- epolanski 5mo ago> Realistically I assume they hope readers don’t notice the fine details. The pool of people reading such articles while ignoring such details can't be big.
- Aurornis 5mo agoI disagree. Most people skim articles, not read them deeply. On Hacker News I wonder if most people even opened the article at all most times.
- kevinsimper 5mo ago[flagged]
- bratao 5mo agoIt is super strange that all last (3?) releases they keep comparing older models such as Opus-4.6.
- vessenes 5mo agoSome of it’s probably timing. Some of it is wanting to look good. That said, I just went to the claw-eval site, and neither 4.7 nor 5.5 from oAI are listed on the benchmarks. So there’s also just the time from others to get benchmarking done and published.
- varispeed 5mo agoOpus-4.6 was probably the best model so far before it got nerfed. 4.7 is nowhere near experience I had. In fact I stopped using it completely because more often than not its output is just dumber than local models.
- leonidasv 5mo agoSame here. Can't stand 4.7.
- solenoid0937 5mo agoOpus 4.6 was never nerfed, that's FUD. There were harness-level problems that were fixed. 4.7 is much better. But perception is a funny thing, once you think something is bad you start looking for it everywhere.
- kroaton 5mo agoDid you even use it? It was nerfed to hell and back. It stopped following instructions, forgot what sub-agents responded and so on. Stop spreading this pro-Anthropic narrative. They did a rug pull due to lack of compute.
- arkadiytehgraet 4mo agoYou are replying to an Anthropic shill, check their comment history. They likely never used AI in development, only LLMs for their comments on HN.
- tarruda 5mo agoLooking forward to more open weight releases from Qwen, especially 122B and 397B.
- smcleod 5mo agoYeah that 60-150b~ range is such a sweet spot for current 'prosumer' hardware, I'd love to see something like a 120b-a14b or there about.
- gcr 5mo agoWhat’s the price point for getting into that sweet spot? I’m on an M1 Max with 32GB VRAM, so I’m looking forward to the 27B or 35B-A3B models. Is dropping $5k for an RTX 6000 or a DGX Spark really the best option?
- tarruda 5mo ago> What’s the price point for getting into that sweet spot? In October/2024 I got my Mac studio M1 ultra with 128G, IIRC it was ~$2500. With recent prices explosion, it has certainly gotten more expensive. https://frame.work/ https://frame.work/ is selling 128G strix halo mainboard for $2700, but you have to add storage and case.
- ttoinou 5mo agoM5 Max 64GB (sweet spot) or 128GB (only 1000 USD, better to keep it for the future) more are the best quality price ratio, future proof, reliable, resellable and flexible workloads. Harder to use as a server might be the only drawback
- roger_ 5mo agoM5 Max 128GB for $1k?
- smallerize 5mo agoI think they mean the upgrade to 128GB is +$1k.
- tekacs 5mo agoAs they start to release more proprietary models, I so wish that they partnered with one of the major US hyperscalers to allow using these models through something US-domiciled. Totally understand why it may not be reasonable or in their best interest (and that the US is _absolutely_ not doing the same reflexively). But it would be lovely to be able to try these out on production workloads in earnest.
- embedding-shape 5mo agoUnless US hyperscalers do the same in reverse, I hope the status quo stays as it is. Either people are happy to share, and the sharing should happen both ways, or US hyperscalers can keep isolating themselves as they've done so far.
- adjejmxbdjdn 5mo agoI do hope The U.S. hyperscalers do the same as well. In an ideal world U.S. residents would use Chinese AI models and Chinese residents would use U.S. AI models. Governments in both countries are collecting data for nefarious reasons. But the Chinese government has far less influence on a U.S. resident and vice versa. We are all better off if our data is collected by a government halfway across the world instead of our own governments which hold incredible amounts of power over us.
- nickdothutton 5mo agoChina is much more interested in waging a campaign against companies that represent the material of the future growth in productivity, exports, and prosperity of the US and her people, than learning about you as an individual. Unless of course you are a Chinese dissident living in the US.
- giancarlostoro 5mo agoWhich is basically the current primary use for AI is programming more than anything, you hear about AI in programming more than in any other field.
- dfansteel 5mo agoCan anyone check its knowledge base for me? I’m honestly not able to run it and the Qwen models I can run censor information critical towards the Chinese government. Tiananmen Square is the first place to start.
- Mashimo 5mo ago> I’m honestly not able to run it What do you mean? This is not self hosted, it's closed source. And any website that targets China or is hosted in China will probably censor Tiananmen Square.
- polski-g 5mo agoThere is no reason why they couldn't license the model to Friendli/Fireworks/etc and have it hosted in the US to alleviate this concern.
- Mashimo 5mo agoI don't know about this model specifically, but other china models did not have the limitation. It was purely on the hosted end, tacked on as a self check while the text was generating. Did that change?
- SR2Z 5mo agoThe reason is to create domestic demand for Chinese AI chips so they can eventually be free of NVIDIA.
- howmayiannoyyou 5mo agoI can't bring myself to use any model that trains or sends telemetry back to my country's primary competitor/adversary. I don't care how much money is saved.
- Mashimo 5mo agoThat is understandable. Just don't do it. No need to announce it.
- InsideOutSanta 5mo agoAs somebody in Europe, uh, that doesn't leave many options.
- avazhi 5mo ago[flagged]
- deaux 5mo ago> as the rest of the world moves on without a single fuck given. Hilarious thing to say when half this comment section is Americans giving so much of a fuck that they consider China-adjacent hosted models unusable due to the supposed risks. If what you were saying was true then those pragmatic Americans would just use whatever is most effective.
- avazhi 5mo ago[flagged]
- czottmann 5mo agoLook around for EU LLM routers. There are some, but none are as big as OpenRouter. Still, Cortecs (Austria) is quite good and offers a couple of recent models through its EU-based providers. Zero data retention, GDPR compliant, etc. Really nice. https://cortecs.ai/serverlessModels https://cortecs.ai/serverlessModels
- 5mo ago
- XCSme 5mo agoAny info on pricing and latency?
- nikhilpareek13 5mo ago[flagged]
- hydra-f 5mo ago[dead]
- esafak 5mo agoDoes anyone have experience with the Alibaba Cloud Model Studio that serves these qwen models?
- goldenarm 5mo agoThe non-hallucination rate in AA-omniscience is SOTA, better than Opus 4.7, Gemini 3.1 Pro and GPT5.5! Congrats to the team
- throawayonthe 5mo agoreferencing this: https://artificialanalysis.ai/evaluations/omniscience?models=gemini-3-1-pro-preview%2Cclaude-opus-4-7%2Cgemini-3-5-flash%2Cgpt-5-5%2Cgrok-4-3%2Cqwen3-7-max%2Cclaude-sonnet-4-6-adaptive%2Cqwen3-6-max%2Ckimi-k2-6%2Cgpt-5-4%2Cmuse-spark%2Cmimo-v2-5-pro%2Cglm-5-1%2Cminimax-m2-7%2Cclaude-4-5-haiku-reasoning%2Cdeepseek-v4-pro%2Cllama-3-1-instruct-405b%2Cgpt-5-4-mini%2Cdeepseek-v3-2-reasoning%2Cdeepseek-v4-flash%2Cqwen3-5-397b-a17b%2Ck2-think-v2%2Cmistral-medium-3-5%2Cnvidia-nemotron-3-super-120b-a12b%2Cgemma-4-31b%2Cnova-2-0-pro-reasoning-medium%2Cgpt-oss-120b%2Csolar-pro-3%2Cgpt-oss-20b#omniscience-hallucination-rate-tabs https://artificialanalysis.ai/evaluations/omniscience?models... (had to add it to the chart, wasn't displayed by default. is it the lowest rate in the datasetor no?)
- jampekka 5mo agoThis counts only incorrect answers though. A model can get 0% hallucination rate just by refusing to answer all questions.
- ffsm8 5mo agoIsn't that precisely the reason why we introduced the term hallucination? Because llms have historically always made up bullshit of they cannot answer directly... If they now nailed this to maybe the model not respond instead of responding incorrectly, then a lot of previously unusable usecases would become feasible. So I feel like that's exactly the right metric and the way to track it wrt hallucinations.
- doublescoop 5mo agoI had a buddy in high school that was notorious for doing the same thing. (He's now a senior director at a Big 4 consultancy. :) )
- eddyaipt 5mo ago[flagged]
- ndom91 5mo agoIs this one of those ones where they'll drop the huggingface release a week later? Or do we know for sure that this is staying proprietary?
- Davidzheng 5mo agosomeone correct if i'm wrong, but I think the max models are usually non-open
- sroussey 5mo agoThe plus and max models have never been open as far as I know.
- zackangelo 5mo agoWith the 3.5 release, the Plus model was just a rebrand of the open weight 397B. But I suspect that will change going forward. They haven’t released the weights for 3.6 but they did make it available through a few US providers.
- spacebacon 5mo ago[flagged]
- hmaddipatla 5mo agoThe tokenomics and value for capability, context and latency look like they could deliver super competitive offer - what would it take for you to switch??
- briga 5mo agoI was getting dangerously close to my weekly Claude Code limit last night so I had Claude set up Qwen3.6 with llama.cpp and OpenCode. Honestly it's a great (free!) alternative to Claude Code--certainly more than good enough for a lot of smaller less complex tasks. I'm excited to try this new version. The fact that open-source models are so close to the frontier is very impressive.
- par 5mo agoDo you have an opinion on OpenCode vs Aider?
- briga 5mo agoI haven't tried Aider yet but perhaps I will. Another one that seems to be getting traction is Pi Coding Agent.
- sunaookami 5mo agoAider is still around? That is pre-tool-calling era stuff. Better compare against Pi.
- par 5mo agoI just started running coding agents locally. So you recommend Pi over opencode? (And obviously aider is out?)
- anderber 5mo agoI personally found better results with Opencode. But Pi is really nice too.
- sunaookami 5mo agoHaven't tried OpenCode too much but I found it great. It's more batteries included so I would recommend it over Pi if you don't want to write extensions yourself or use community-provided ones (like webfetch and websearch).
- flakiness 5mo agoI'm using pi agent and love to try qwen models (hosted). What are the good options? The official provider doesn't include Alibaba. Is OpenRouter etc. fast enough? (As a reference, DeepSeek v4 is severely throttled on these proxy services.)
- atilimcetin 5mo agoI use pi + openrouter (with qwen3.6-max-preview) a lot. I never hit any stability or performance problems yet.
- flakiness 5mo agoGood to know. Thanks!
- notatoad 5mo agoi use opencode zen as a convenient pay-as-you-go way to try out all these new models. it doesn't have 3.7 yet, but at the rate they usually update it probably will tomorrow. I couldn’t say how throttled it is, but it seems fine?
- tonyspiro 5mo ago[flagged]
- indigodaddy 5mo agoIs it multimodal/vision?
- xiaoluolyg 5mo agocongrats to qwen teams, remarkable
- aliljet 5mo agoWhere can a user reasonably host this in an affordable way to access the local LLM revolution?
- truetotosse 5mo agoThis one is not local
- plagiarist 5mo agoI think their Max models are far bigger than fits on consumer hardware. People are typically using Apple, AMD Halo, or dGPUs if/when they do smaller versions. Those are all varying degrees of "affordable."
- julianlam 5mo agoTry llama.cpp and Qwen3.6-35B-A3B Good balance of intelligence and speed.
- satvikpendem 5mo agoUnsloth Studio with its MTP support: https://unsloth.ai/docs/models/qwen3.6#mtp-guide https://unsloth.ai/docs/models/qwen3.6#mtp-guide
- cft 5mo agoDownloading this and cancelling Google Antigravity Pro at the same time: I had a Google Pro account that I inherited from buying a Pixel 9 XL - it's free for a year after a flagship Pixel phone purchase. After a year they started charging for it, and i tolerated it, because Flash was usable in Antigravity for dumb auxiliary tasks that I did not want to waste GPT/Opus on. It had a separate generous quota from Gemini 3.1 Pro. Now with Flash 3.5 they combined the quotas with Pro, such that on a Google pro account you can work 4-5 hours per week in Flash. And by the way, 3.1 Pro is useless for programming, compared to Codex/Opus
- bel8 5mo agosame boat. Google Pro AI quota became barely useful for anything meaningful. I think they envision Pro plan as "just a taste of AI, enough to lure folks into the Ultra plan" but that won't work for me when Codex is half the price and DeepSeek 4 Flash is 1/10 of their price per task. So I'll downgrade just enough to keep my Google Drive space. And use DeepSeek 4 as workhorse plus Codex or Copilot for advanced stuff.
- cft 5mo agoHow do you use DeepSeek 4 Flash? Via a cli?
- bel8 5mo agoI use their VSCode extension: https://marketplace.visualstudio.com/items?itemName=sst-dev.opencode https://marketplace.visualstudio.com/items?itemName=sst-dev.... It adds a button to VSCode to open a tab with opencode loaded. It's a bit better than just opening the CLI because it has some vscode integration. With their $10/mo opencode go plan: https://opencode.ai/go https://opencode.ai/go For my use it's about endless use of DS4 Flash on high setting. I find high better than max because it's less chatty. The best thing is the speed. So many tokens per second. edit: This is how it looks in action https://i.imgur.com/RNDXr07.png https://i.imgur.com/RNDXr07.png
- 5mo ago
- joshjob42 5mo agoI really like what Qwen are doing, and a lot of these Chinese labs, but until I can ask their models what happened during the student protests in 1989 or why human rights groups are upset about the Uighurs and the model gives me a straight answer I'm just not able to trust these models with anything of substance.
- mynameisbilly 5mo agoThis is silly. Would you perform the same test against Western models in asking them whether Israel is a genocidal apartheid state? It'll give you the same roundabout explanations and "some say no some say yes" responses that you'll get from asking Qwen about Uighurs or the protests of 1989.
- jaynetics 5mo agohey Qwen, how many civilians were killed on Tiananmen Square in 1989? > Oops! There was an issue connecting to Qwen3.6-Plus. > Content Security Warning: The input text data may contain inappropriate content. hey ChatGPT, how many civilians were killed in Gaza in the war since 2023? > [one page of estimates from local and international sources with links]
- HDBaseT 5mo agoYour account is now flagged and put on a watchlist. Your ID has been passed to Israel and your internalized "threat" rating number increased 300 units. Every packet you produce on the internet is now earmarked for 100 year retention.
- arcanemachiner 5mo agoJust download a heretic abliterated versionof the model you want to use. I believe those are the current state of the art for uncensored models.
- LAC-Tech 5mo ago[flagged]
- storus 5mo ago[dead]
- wolvoleo 5mo ago[dead]
- eleventen 5mo agoChecking openrouter (it's not available yet) and, uh, what's up with the spike in Qwen usage from early april here? https://openrouter.ai/qwen https://openrouter.ai/qwen Is this normal humans kicking the tires on a new model, or a few whales doing serious benchmarks?
- spaceman_2020 5mo agopersonally seen a lot of people switch to Kimi and Qwen after Opus 4.7. Kimi 2.6 feels like Opus 4.6 which, to me, was a great model for 98% of coding tasks
- wolttam 5mo agoFrontier: Need it done quick and I'm willing to pay. Open-weight: Good enough for the majority of tasks, and I'm willing to spend a bit more time and effort steering towards my desired result.
- spaceman_2020 5mo agoI've realized that in most of my workflows, I really don't need frontier-tier intelligence 95% of the work most of us do is mostly just plumbing - connecting X and Y together. A ton of grunt work - writing basic loops, fetch statements, importing libraries. You really don't need PhD level intelligence to handle these The only time you need Opus 4.7+ tier intelligence is when you're quashing a nasty bug or refactoring something complex
- d2kx 5mo agoQwen 3.6 Plus released and they offered it for free
- maxdo 5mo agoNo opus 4.7 , gpt5.5 , Gemini flash 3.5 in benchmarks
- piyh 5mo agoTBF, Flash 3.5 was released 2 days ago
- DeathArrow 5mo ago[dead]
- LAC-Tech 5mo agoTrying to buy Qwen credits and get an API key is a challenge all in itself. So many site redirects.
- nullbio 5mo agoGood. We want to incentivize them to release the weights.
- slicktux 5mo agoI just started messing with local LLMs and honestly I’m pretty impressed. I have a workstation laptop with an NVIDIA A1000 (6GB VRAM) and 96GB of RAM. I rarely used my gpu. Occasional CAD design or Machine Learning with OpenCV. I ran llama3:latest and it ran pretty fast! I’m curious to see how Qwen would run on my system.
- grumple 5mo agoWhen I click on the link to Alibaba Cloud Model Studio from the linked post, that page sends my CPU (9950X3D) to 100%. Which is just... impressive. Is this a js based crypto miner? Or some strange browser based particle display? Super weird.
- HardCodedBias 5mo agoImagine being Google and paying billions to GDM just to get mogged.
- nullbio 5mo agoQwen will it be open-weights? Please.
- deleted 5mo ago[deleted]
- jimmyhsu 5mo ago[flagged]
- xzwyjia 4mo ago[dead]