13 ms·
They've also announced Qwen3.8-27B being released open-weight next week. Qwen3.6-27B is widely regarded as one of the best local models, especially since nothin
by toshinoriyagi 2mo ago
They've also announced Qwen3.8-27B being released open-weight next week. Qwen3.6-27B is widely regarded as one of the best local models, especially since nothing else comes close to it, that isn't benchmaxxed, without being significantly larger. If 3.8 truly improves upon it that would be awesome.
- icelancer 2mo agoThis is what I've been waiting for. We are still using fine-tuned deployments of Qwen3.6-27B with a lot of success but could use a bump in intelligence. Here's hoping.
- magicalwh 2mo agoHow and where do you finetune it?
- leansensei 2mo agoEasy to do with Unsloth Studio.
- icelancer 2mo agoI use Modal for fine tuning and unsloth mostly.
- nozzlegear 2mo agoQwen3.6-35B is my daily driver for AI, and what convinced me to cancel my Claude subscription back in April. The Qwen3.6 line is easily the best local model I've tried, and I've tried a lot. I've got it diligently grinding away on my laptop right now, reviewing and fixing some bugs in my F# code.
- pettijohn 2mo ago35B MoE is certainly a good and fast local model. I find 27B dense to be quite a bit smarter, so I daily drive that. I wish there was a ~100B MoE with maybe 10B active. It would be super smart and fast!
- nozzlegear 2mo agoI've heard 27B is smarter! I tried it some time ago but couldn't get it working with my oMLX. I need to try it again.
- mattnewton 2mo agoHonestly the 27b dense one punches way above its weight in a lot of domains, especially coding in my testing, so I think you will probably be disappointed.
- npodbielski 2mo agoIn my case I would say they are comparable but moe models are looping and getting lost a lot more than dense models. On the other hand having 90t/s with any local model is nice and Pi with loop police extension can prevent looping a lot.
- cyanydeez 2mo agoLooping seems related to quantization and not the model itself. If youre digging deep into quants to get working context then yeah.
- mattnewton 2mo agoThere was a 3.5 122B 10A release - https://huggingface.co/Qwen/Qwen3.5-122B-A10B https://huggingface.co/Qwen/Qwen3.5-122B-A10B
- kanemcgrath 2mo agoI tried it for a bit, and It was not really worth its size. It got swept up in all the other AI news recently, but laguna s 2.1 I think is the best ~100B moe model right now
- nozzlegear 2mo agoI didn't mention it above, but Laguna S is my other favorite model. I use Qwen a lot more, it's smaller and faster, but I like to switch to Laguna when I feel like I need a "heavy hitter" for certain huge or complex tasks.
- neumann 2mo agocompared to claude - how 'fast' is it in terms of throughput on your laptop?
- syntaxing 2mo agoI use it with a strix halo server. 35B runs stupidly fast. 27B is about 700 TPS prefill and 30 TPS token generation. Which interestedly is about what Kimi K3 gives me depending on provider.
- dionian 2mo agowhat hardware do you use or recommend for this? never heard of it until today.
- Zetaphor 2mo agoStrix Halo is the unified memory platform from AMD. Similar to the DGX Spark from NVIDIA or the M series Macs. I personally have the Framework Desktop, but there's also systems from other brands like Bosgame
- vrganj 2mo agoYou can also get it in a laptop form factor that feels like a MBP with a nicer keyboard if you get an HP Zbook G1A! Huge fan of that thing, it's th e Linux MBP I've always wanted.
- Zetaphor 2mo agoWhile the laptop option is nice, for an inference server you're probably going to want the desktop form factor as it has significantly more thermal overhead and thus better performance. In the desktop models most of the internal volume is a gigantic heatsink
- leansensei 2mo agoRTX 4060 and above. Ideally RTX 50 Series, because you can run NVFP4-quantized GGUFs that give you better prefill AND better quality.
- ufish235 2mo agoWhat laptop?
- nozzlegear 2mo agoIt's just a MacBook Air with an M4, cheap and nothing special. I host Qwen on my Mac Studio, an M1 with 64gb ram. The model uses around 20-25gb ram depending on what it's doing.
- brailsafe 2mo agoI wonder if I could get this running on my 48gb M4 Pro. Haven't been able to load anything beyond 27B
- deleted 2mo ago[deleted]
- deleted 2mo ago[deleted]
- razster 2mo agoI would recommend looking into Ornith1.0 - it's using Qwen3.6 35B-A3B and excels in coding, at least for my coding needs, Python, web-dev, SQL scripting and some C#. Using Pi harness.
- mraza007 2mo agoI have been using Qwen3.6-35B-A3B as my daily driver as well and its been phenomenal when it comes to coding
- coverband 2mo agoHow do you use a 72GB model as your daily driver locally?
- chmod775 2mo agoAt 4 bit it easily fits in 32GB. That's what most people use.
- sznio 2mo agouse a quantized version. since it's MoE, what matters is that the 3b parameters that are used for every token fit in gpu vram, the rest can stay in system ram. really great if you don't have unified memory.
- anon373839 2mo agoApple Silicon. But: there's no need to use the FP16 version. At 8-bit precision the quality loss is almost imperceptible. That cuts the footprint to 36GB. Which is great for a 64GB Mac, because you have room for plenty of context. 6-bit also works nicely at 26GB + context. You want to use the newer quantization formats like Unsloth's UD quants or oQe, where the weights are selectively quantized using a calibration dataset so that important weights are left at/closer to full precision.
- trollbridge 2mo agoQwen-3.6-35B-A3B was our "gateway drug" into switching our organisation to agent/harness-first coding. Particularly, I had one team member who was extremely sceptical of AIs/LLMs/harnesses and refused to use them. One day he said "Well, I have an RTX 5090 doing nothing... should I try to get something up on it?" and a few minutes later he had 3.6-35B loaded up, running OpenCode. It continues to be a workhorse to this day, running on both my local Mac for various types of jobs, an AMD R9700 at the office, and said teammember still uses it on his 5090, although in practical terms we do a lot more with DS-V4-Flash-0731 these days.
- websap 2mo ago[flagged]
- badsectoracula 2mo agoA local model needs 0 investment and 0 commitment, takes literal minutes to get started (especially if you have someone who is into that stuff showing you the ropes) and if you end up disliking the experience of using AI you can just `rm -fr` it and forget the whole thing existed.
- jiggawatts 2mo agoThis is the diametric opposite of the rent-vs-buy scenario that this entails. Local: You need to invest $thousands into GPU and/or very-high-end CPU+Memory hardware. Vendor: You can use any existing device, even a phone or tablet. A very low-end laptop is fine. > takes literal minutes to get started Local: Typical scenario is hours just to download the software, the model weights, and then faffing around with CUDA and matching your GPU drivers. Vendor: Free-tier available instantly on a web URL. Even local agents have free tiers from multiple vendors. Install is a single command and/or download and "next,next,next,finish" wizard that takes ~1 minute. > you can just `rm -fr` it and forget the whole thing existed. I'm still cleaning up multi-GB model weights floating around in hidden subdirectories under my user profile from months ago when I was experimenting with local models! Meanwhile I simply... stopped using Gemini. That was the entire process: I no longer actively use it. They stopped billing me for my token usage, because it is now zero. That's... it. You have it totally backwards.
- bckr 2mo agoWhat are the specs of your laptop and what tokens per second do you get?
- nozzlegear 2mo agoIt's just a Macbook Air with the base M4 and 16gb ram, but I'm hosting the models on a Mac Studio with M1 Ultra and 64gb ram that I had purchased when it came out. I get about 45-55 tokens per second with this setup. I think I could get more if I spent some time fiddling with the parameters, but I don't really know what I'm doing there so I've just left most of it on oMLX's defaults.
- apexalpha 2mo agoIt was for me too but the new deepseek pricing is too good to ignore for now. I honestly think that with my electricity prices running qwen 36B myself is more expensive than hitting the cache rate at deepseek.
- sparkling 2mo agoCan you elaborate on DeepSeek (deepseek-v4-flash, i assume?). What does your typical usage pattern look like and what is your weekly/monthly spend? I gave it a try for a few days (pi + openrouter + deepseek-v4-flash via deepinfra) and ended up paying ~$18 for rather light usage. Yes it's still cheap, yes it's fast, but i feel i would still get a better deal with a Claude subscription plan.
- maattdd 2mo agoDeepSeek without OpenRouter is wayyyy cheaper
- sparkling 2mo agoIs it? OpenRouter shows DeepInfra being cheaper than DeepSeek directly https://openrouter.ai/deepseek/deepseek-v4-flash-20260731#providers https://openrouter.ai/deepseek/deepseek-v4-flash-20260731#pr...
- maattdd 2mo agoInteresting. My anecdotal experience is that it is.
- thecopy 2mo agoI agree with parent. OpenRouter might be cheaper list-price, but i have been using 10$ on DS platform since April/May, still have 2$ left. Using OpenRouter i depleted the same dollar-amount in a 1-2 weeks with same usage pattern. No idea why.
- 2mo ago
- westpfelia 2mo agoI'm a big ole noob when it comes to local AI. What are you using for a harness? Or platform to interact with it?
- ch_sm 2mo agoyou can try ollama, omlx or llama.cpp for instance to download a model and get an inference server running locally. They expose „open ai compatible“ endpoints, so you can configure almost any harness to use them.
- jlkuester7 2mo agoI recommend trying pi.dev as your agent harness for local models. In my experience it has been the sweet spot of functionality (which you can and should extend with plugins) vs performance (OpenCode just swamps local models on my hardware).
- kumarvvr 2mo agoOn what hardware do you run the model locally, if so?
- anon373839 2mo agoNot the GP, but I run this model as daily driver too. It runs great on a Macbook Pro 64GB (M3 Max). Token generation speed can be about 100 tokens/sec with multi-token prediction, although it depends on the context. Worst case speed is around 50 tokens/sec. The weaker point is prompt prefill, which starts at 1,400 tokens/sec but decreases significantly at high contexts. That said, for agentic scenarios, if you're using a harness that doesn't needlessly bust the cache, it doesn't feel slow. I really hope they release a Qwen 3.8 35B, although the lack of a mention seems ominous.
- cicko 2mo agoWith laptop being ...?
- nozzlegear 2mo agoIt's just an M4 MacBook Air with 16gb ram – probably incapable of running models itself. I actually run the models on my Mac Studio which is an M1 Ultra with 64gb, and oh-my-pi on my laptop is configured to use the models over the local network.
- wookmaster 2mo agoYeah I have 24 GB Ram on my M4 and am disappointed in the capabilities. You really need a ton of RAM. I can run some basic 8B models fine but they dont do well at all in basic coding stuff I've thrown at them.
- quanto 2mo agoIs your Qwen3.6 locally run on your laptop? What kind of tokens/s are you getting from your laptop GPU?
- saagarjha 2mo agoI have Qwen3.6 35B-A3B on my laptop and it does 60 tokens/s
- nozzlegear 2mo agoQwen is running on my Mac Studio, an M1 Ultra 64gb. My harness (oh-my-pi) on my laptop is configured to use the models hosted on my local network, since it's just a MacBook Air 16gb and probably incapable of running anything useful itself. I get about 45-55 tokens per second using Qwen with this setup. I could probably squeeze out more if I messed around with the settings, but I'm mostly using oMLX's defaults for the model.
- btbuildem 2mo agoI'm on the verge over here, the new Anthropic models have been a disappointment. I've tried the A3B variant, but had mixed results. What do you use as the coding agent, and have you heavily customized your workflows?
- ncphillips 2mo agoWhat kind of machine do you have running that? My attempts at local have always resulted in a very hot lap
- AbsurdCensor 2mo agoStrix Halo for me. If I am running something on my laptop, it's a much smaller usually around 12b model, but those are a bit less functional. I mean I think there is a a ROG FLow Z that has the Strix Halo setup, but that thing was super expensive.
- nozzlegear 2mo agoI host the models on my Mac Studio, an M1 Ultra with 64gb ram (I bought it when it came out, just happens to be good at LLMs). So when I work on my laptop, I have my oh-my-pi setup configured to use the models on my Mac over my local "bonjour" network or whatever Apple calls it. That way I have a nice cool lap, while using models that my M4 MacBook Air with its 16gb ram couldn't possibly run.
- ncphillips 2mo agoCool yeah. I got a M3 with 32GB ram and it’s a little iffy. I’ve considered getting a MacMini to act as an in-house
- bibstha 2mo agoWhat do you use to pair it with web search?
- AbsurdCensor 2mo agoDepends on how you are doing it. LM Studio and tool calling models can use the web, or you could go for something like Perplexica, or if you want to go real crazy, something like Hermes or OpenClaw.
- nickthegreek 2mo agoI use searxng.
- bmitc 2mo agoHow do you run Qwen? I tried setting it up the other day and got really confused with all the options between Ollama, LM Studio, and llama.cpp and also the bazillion different Qwen models available.
- alasr 2mo ago- See "Llama.cpp vs Ollama"[0] and "Llama.cpp vs LM Studio"[1] for a high-level features comparison and find out which one suite you best (check out the "Target Users" category[0][1], at least). - If you're still not sure, "LM Studio"[2] is ok to start with as you'll be able to download, start/stop/manage and chat with your LLM model all in one place: a single desktop app. Also, once installed, enable "Developer mode" under "Settings > Developer" tab; you might find it useful later. - Regarding LLM models you can, based on your hardware spec, start with Qwen3.6, Google's Gemma4, OpenAI's gpt-oss-20b or Nvidia's Nemotron; use LM Studio's "Model Browser" screen to search for LLM models (each one listed with their organizational name/brand & logo: ignore the one you don't recognize as you might not need them at the start of your journey; you can always revisit them later, if needed.) - Regarding which quantized LLM models you should download & run, just go with the default "LM Studio" selection (at least, in the beginning). Later, you can experiment with other quantization values to find out which one works best for your use cases. I hope it helps. --- [0]. https://llama-cpp.com/llama-cpp-vs-ollama/ https://llama-cpp.com/llama-cpp-vs-ollama/ [1]. https://llama-cpp.com/llama-cpp-vs-lm-studio/ https://llama-cpp.com/llama-cpp-vs-lm-studio/ [2]. Ignore "LM Studio Bionic" for now; just download and install "LM Studio" from https://lmstudio.ai/download https://lmstudio.ai/download
- bmitc 2mo agoThanks! Unfortunately, LM Studio does not support network proxies, and Ollama does not support Qwen3. So it seems like Llama.cpp is the only solution.
- XCSme 2mo agoIf they trained it well, and can do computer use, it will be a new era. Companies can keep PCs, put Qwen 3.8 27b on it and get rid of the employees, lol...
- deleted 2mo ago[deleted]
- hippycruncher22 2mo agoYes let’s get rid of employees so no one is employed but somehow they can afford to buy my stuff
- overfeed 2mo agoIt's the natural outcome of next-quarter short-termism. The board and C-suite will be fine (monetarily).
- XCSme 2mo agoI am surprised that they keep going with it, seeing how fast it improves and basically soon running themselves too out of business. What's even their end goal? Open source models make sense, if profit is not the target, but for OpenAI and the rest, once they achieve "AGI", don't they basically become useless?
- asqueella 2mo agoNot a novel idea! https://xcancel.com/jurijkovalenok1/status/1863243462113398922 https://xcancel.com/jurijkovalenok1/status/18632434621133989... (Herluf Bidstrup's "Automation")
- mathieudombrock 2mo agoQwen 3.6 27b has been the sweet spot for me in terms of local models. I've had good luck using it with Pi harness. Looking forward to this.
- cybertim 2mo agoI also "evolved" into 27b (q8 unsloth) and pi.dev (tried many combinations) feels for me the same as opus 4.5 that i use at work, faster even (using 2x 3080 20GB gives me 60-80tk/s). Though you do need to feed it more details up front (about what exactly you are planning to do and a good written skill.md) but I work that way anyways, im hyped for 3.8
- iagooar 2mo agoHaving invested in a machine with 128GB of RAM, I would love seeing something a bit larger than 27B / 35B, possibly a 54B dense model or 70B MoE would be much closer to the Qwen 3.8 Max experience.
- hadlock 2mo agoAll of us with a 96gb rtx 6000 would love to see a 70b moe. Maybe they are waiting for OpenAI and Anthropic to IPO so they can short their stock and release. Local LLM is going to get very interesting in the next 2 years.
- ksec 2mo agoFor those of us who don't have the time to follow closely, Qwen3.6-27B being Open Source and Open Weight, what level is this compared to other Western paid version? Just so that we know what 3.8 would be like. I currently have about 150 Tabs of Antirez posting on AI and running local model I haven't had the time to read. And there are probably some prerequisite reading or other research in between as well. I just wish there are some very high level overview and news coverage on all these.
- wickedsight 2mo ago> what level is this compared to other Western paid version? IMHO this is a difficult question to answer. Part of the power of paid models comes from the software supporting it. With local models, you have tons of workflows that can severely influence the quality of the result. In my personal experience, the SOTA models are way more consistent and can handle more complex questions. Part of that is (probably) because I don't let my local model access the internet, while paid models do use the internet to look at docs etc.
- InsideOutSanta 2mo agoYou absolutely need to let models access the Internet if you want consistently good results. Pretty much any non-trivial task requires the model to do things like look up APIs, code examples, or existing discussions of a given topic.
- ndriscoll 2mo agoYou don't really need the Internet. Tons of documentation is available for download (either as a zip, or with the documentation site as its own git repo). Wikipedia is available for download. You can get reddit dumps, HN dumps, stack exchange dumps, etc. This can all easily fit on one hard drive. Reading the actual code is also always a better source of truth than docs anyway (this is true for people and LLMs). Just clone whatever libraries you use.
- c16 2mo agoSome advice I got from another HN Mac user was to run local models in energy saver mode. You'll get slightly reduced tokens, but the laptop won't overheat and the fans won't go wild.
- fragmede 2mo agoOh. I've been using an icepack under my laptop to keep mine cool. I'm watching it with llamatop to see if the GPU is actually active or not, aw activity monitor wasn't showing me what I wanted.
- hooli_gan 2mo agoThat's exactly how i fried a laptop. The condensation killed it.
- leansensei 2mo agoAnd if you're running it on a dGPU, power limit it, because you lose very little in terms of token generation performance, since it's memory-bound.
- throwaw12 2mo agoQwen3.8-Max is the first in Qwen-Max series to be open-weight as well. Kimi K3, GLM 5.2 and now Qwen3.8-Max - open weight models. DeepSeek V4 Flash outperforming Gemini 3.1 pro, probably DeepSeek V4 Pro update is also coming soon Chinese labs are cooking very hard. US closed weight labs are probably hard time to resist not calling Washington DC for more AI regulations
- badsectoracula 2mo agoKimi K3 is more like "weights available" in that you can download and use them but it is under a custom license that has a bunch of limitations where you have to pay Moonshot for doing some stuff. GLM 5.2 on the other hand is plain old MIT. Not sure how Qwen3.8-Max is going to be licensed, hopefully it'll be Apache like the smaller ones.
- pama 2mo agoYou can do whatever you want with the model within your own organization. If you use it commercially—either as a model-as-a-service business or in a very large-scale product—you should check the additional license terms, which go beyond MIT. My interpretation is that Moonshot cares about the exact inference behavior and accurate representation of their model or derivatives, and perhaps also about capturing some additional value despite their own GPU limitations, so the extra license terms focus on those large-scale commercial deployments.
- yassa9 2mo agoas much as im excited for it, sadly it gonna be one of the reasons to push ram prices higher
- foft 2mo agoReally awesome. Though I wish they'd do a dense 48B, 60B or 72B. There seems to be quite a gap between the small ones and the enormous ones these days.
- bertili 2mo agoThe 27B have many more active parameters than much bigger models such as DS4Flash, MiniMax etc, which makes it punch above its tiny weight. A great fit for a 5090 in a closet for meat-and-potatoes, kind of work.
- SillyUsername 2mo agoDo you have a news source for 3.8 27B pls?
- marci 2mo agohttps://xcancel.com/Alibaba_Qwen/status/2084100707423289643#m https://xcancel.com/Alibaba_Qwen/status/2084100707423289643#...
- LeBit 2mo agoYou consider that even Gemma 4 31B is not even competing with Qwen 3.6 27B?
- HarHarVeryFunny 2mo agoThere was an interesting interview by MLST with a team doing well on ARC AGI 3 who are using Qwen 3.6 27B, and said that it's actually better at coding than the larger 3.6 35B. I guess which of the smaller 3.8 models is best for coding will depend on which one they put the training effort into.
- als0 2mo agoThe larger 3.6 35B is actually a mixture of experts (MOE). This means a small proportion of those B's are actually active. It's fast and suitable for agentic tasks but nowhere near good as the dense 27B model, which has all of its parameters loaded.
- colingauvin 2mo agoI have found 27 to just be so much more coherent than 35: https://humanparadox.org/local-vs-frontier-benchmarks-for-my-personal-assistant/ https://humanparadox.org/local-vs-frontier-benchmarks-for-my... It can complete multi-step tasks much better, and has a bit more curiosity.
- jedbrooke 2mo agothe bonsai 27B 1bit quant version of Qwen3.6 27B is even more nuts, model fits in 4GB, and with 100k of content model+kv cache fits in 8GB. I’ve been running it locally on my mac mini 16GB. it gets around 4-6 tok/s, so not quite real-time ready, but good enough to let it run on task async for 20 min and come back. The 1bit model struggles a bit with multi-turn conversations though (e.g. when switching from plan to act mode it will still keep trying to make a plan) but that’s easy enough to reformat prompts into multiple one shot sessions of smaller work.
- diseasedyak 2mo agoI've been running Qwen3.6-27B-IQ4 (4-bit quantized) locally and it's been great. I can't run the non-quantized version as I only have a 4090 w/24 GB of VRAM and it won't fit and leave any context room, but the quantized version only uses 18GB.
- aiagenta2z 2mo agoYeah, Qwen3.8-Max is the new Flagship model for coding and harness system and many other benchmarks are reaching equal performance as Claude and other close models. That's gonna drop the price of LLM in agent landscape a lot.
- ttflee 2mo agoIt is really a difficult choice in front of me: DeepSeek v4 flash 0731@q2 vs Qwen 3.{6,8} 27B@fp8 on a 96GB VRAM server.