14 ms·
Qwen 3.6 27B is the sweet spot for local development
- 217 3mo agoThis is kind of like saying grass is green to be honest
- madduci 3mo agoLike everybody got 128 GB RAM..
- dofm 3mo agoDoesn't need it at Q4 at least; it'll run in 64GB.
- intothemild 3mo agoQ6 can run with 256k at Q4 on 32gb easy. 200k @ K : Q5_0 V: 4_1 (which is a bit of a sweet spot)
- sleepyeldrazi 3mo agoI've been running it almost since launch on a 3090 (24gb vram), you really don't need that much. Second hand those are really cheap and i get 50-70 t/s (with MTP at 2), full ctx. IQ4_NL (unsloth) on this model seems suspiciously competent, and after the (by now not so recent) updates to q4 KV on llama.cpp, I just keep going back to it after dsv4pro disappointed me for the 100th time because it gave up on a task.
- aand16 3mo agoI've come from the future to say Qwen 3.7 27B is just around the corner and slaps!
- lor_louis 3mo agoDo no give me hope like that.
- mendeza 3mo agoI am eagerly waiting!
- jensC 3mo agoMe too, I am on a Jetson Orion 64GB (about 50W max). Using the nvidia graphic cards for AI seem to be so power hungry that it was not a choice I could take with todays environmental problems.
- NamlchakKhandro 3mo agoHuh?
- layer8 3mo agoAre RAM prices down?
- aand16 3mo agoNever did. It was a while back, but in the great Chips War the Micron and Hynix fabs got nuked to atoms. So all memory was confiscated for the war effort. But now that things are finally turning normal, everybody gets a daily token allowance for GeminiGPT 7.2 and you can even boost your ratio if you volunteer your bioenergy to the local GPT microdatacenter (cycling or run on the H-wheel for an hour gets x1.02 output tokens for the day).
- alfiedotwtf 3mo agoQwen 3.7 120B will kill off Antropic’s IPO
- HotGarbage 3mo agoAnd AI companies will continue to buy up all the silicon to make this prohibitively expensive to run at home.
- dofm 3mo agoIt will run (somewhat slowly) on a five year old M1 Max with 64GB RAM. Personally I prefer the 35B MoE model, which is fast enough to be interactively useful, and capable, but I would probably use the 27B if I wanted to generate whole applications like that. I am unconvinced that most "local" AI applications need anything much more powerful than the Gemma 4 12B model. Local agentic coding is a small niche, but there are plenty of ways a local model can help with development tasks. I would really like to see a 12B or 16B Qwen 3.6. I am currently playing with Ornith 1.0 in the MoE configuration, which is based on the 35B variant of Qwen 3.5; I am not sure if it is better than the 3.6 version. Benchmarks say it is; my own silly tests either suggest otherwise or suggest that I have to talk to it a bit differently.
- sleepyeldrazi 3mo agoI need to ask, since I have desperately wanted to make Gemma 4 12B work, but im not sure if its the quant (i usually up it to q8, which is a lot higher than iq4_nl that i use for 3.6 27B) or the model itself, but it just starts confusing itself really quickly when I give it coding tasks. And quickly starts failing tool calls. I really want to have a model that i can run locally on my 24gb m4 pro mbp for when i don't have internet to connect to my 3090 running the qwen, and i love how gemma 4 models 'feel', but i can't make them be competent. I am in the middle of finetuning both qwen3.5 9B and gemma 4 12B just to try and make those bridge closer to 27B for coding/agentic tasks (and am trying to ternarize and DQT 27B so that it fits in ~9gb pre-KV). How do you run the gemma? What do you use it for (and in what harness), maybe llama.cpp and pi-mono just aren't for this model and that's what i'm doing wrong.
- dofm 3mo agoIt sounds to me like you're further along on this than I am, if you are fine tuning? I am still mostly tinkering/learning rather than spilling out code, and I feel quite slow on it. So it doesn't matter too much to me if it is really slow. More the journey than the destination if that makes sense. I'm stubborn. I have tried the Gemma 4 12B model (Unsloth's QAT version) with search/browse tools in LM Studio and Unsloth Studio, when I am trying to understand a new thing. Basically I get it to write introductory starter documentation for me to absorb, because my big personal problem, these days, is focussing enough to start a project and then digging in; I need the help. I have found its limits on obscure packages (that it sometimes makes up) but before that it's a bit like stumbling on a blog post that happens to be really right for your particular need. Good enough to work through. It's stuff I could ask Perplexity to do, or ChatGPT, to be fair, I just like LM Studio for this and have the inquisitiveness to want to run it locally. In your case: I don't believe it's the quant. I'm sure it's the model — it has good coding knowledge but it's clearly not specialised. It might be good enough at writing Python/PHP/JavaScript at a novice level. It is also quite good on WordPress tooling and functions. But I wouldn't bother with it for agentic coding if you've got experience elsewhere. Might be interesting to see what you can do with the 9B Ornith model? Qwen 3.6 MoE in its Unsloth version is another matter. Impressive and I am trying to find ways to support my old brain doing what I've done before.
- rusk 3mo agoSpent a week trying to get sensible results out of llama 3.3 At one point it even simulated doing the work, log output and everything and when I challenged it about the missing artefacts it actually started questioning my intelligence. Seems appropriate for a Zuck enterprise. Qwen on the other hand got straight to work with astonishing competency on the same system. From what I read llama3 needs beefier compute to reliably invoke tools, which I presume relates to it focussing more on simulating AGI rather than being a useful tool.
- am17an 3mo agollama 3? Are you from 2023?
- culi 3mo agoYou might find this helpful. llama is not anywhere near the Pareto distribution (performance vs cost) https://arena.ai/leaderboard/code/webdev/pareto?license=open-source https://arena.ai/leaderboard/code/webdev/pareto?license=open... https://arena.ai/leaderboard/text/pareto?license=open-source https://arena.ai/leaderboard/text/pareto?license=open-source
- k__ 3mo agoLlama3.1 instruct seems to be doing okay on that page, mostly because it's dirt cheap.
- rhgraysonii 3mo agoI have been having pretty good success with Qwen 3.5 9B for "nontrivial but not challenging work all things considered" -- it runs great on my 24gb unified memory m4 pro MacBook Pro. What do the baseline specs look like Mac-wise for getting this model to run? Am I looking at a 96gb? 128? 256?
- dofm 3mo agoYou might be interested in Ornith 1.0 9B, which is a new intriguing post-training of Qwen 3.5 9B. Qwen 3.6 27B will run in full offload with a 4-bit quantisation in 64GB on an M1 Max. It is quite slow. I don't know about 48GB but 64GB should be enough.
- rhgraysonii 3mo agoThanks! I was thinking of doing the 128gb to have some future proofing. I figure at this point, it's akin to a mechanic keeping great tools around, when it comes to having this sort of homelab and exposing it for your own uses. And great practice for building the next era of user facing computing that will be around as this proliferates.
- dofm 3mo agoI would not buy a 64GB model again, probably, if this were to remain particularly important to me. But I gather memory bandwidth is pretty important here. So for example I'd favour a used M1 Max over a used M2 Pro, at least based on my naïve understanding. Not quite sure where the balance changes. There appear to be some hardware improvements with the M3 and up regarding the Apple Neural Engine which I'd hope would show up in MLX performance; I remember seeing some optimisations in image generation models that are only possible on later hardware. The GPU cores are progressively better I believe, but the memory bandwidth is lower. Though perhaps the M4 can get closer to actually saturating said bandwidth. (And I must reiterate that my understanding of this stuff is pretty naïve.)
- freehorse 3mo ago
- blobbers 3mo agoHow does llama.cpp use the GPU efficiently as opposed to MLX? Is there any way to use MLX and GPU at the same time? Or does memory become a big problem? TBH, I never understood Apple hyping these neural cores because I didn't think anyone actually uses them except maybe certain photo/video editing software. If I can generate voice at the same time as video, that would be useful.
- dannyw 3mo agoLlama.cpp uses the GPU very effectively because inference of LLMs is very rudimentary and basically as simple as your GPU memory bandwidth. That's essentially the baseline performance ceiling, with model-specific optimisations like MTP potentially increasing it. The neural cores aren't suitable for LLMs/transformers and isn't used in LLM inference. On the M5 and later chips, it comes with neural accelerators, aka Tensor Cores, which speed up the 'prefill' (i.e. processing your context window) part, but don't do anything for inference. The MLX vs GGUF debate is mostly irrelevant. The GGUF pathways are optimised for apple silicon to the extent of practically identical performance to MLX. MLX is just one way of using Apple GPUs, it comes with many optimisations in the box, but they're not hard and they're no longer MLX-exclusive.
- kpw94 3mo ago> What it does: > > --jinja for tool calling support Pretty sure this flag hasn't done anything for a while. It's enabled by default since ~November of last year
- ascii0eks84 3mo agoVery capable lora adapters are surfacing but it seems they are very niche.
- DenisM 3mo agoCan you share more? It’s the first I hear of lora outside research papers. Practical applications would be great to see. Lora if effective could be a great reason to run local models.
- 0x0000000 3mo ago> ... on my Macbook Max M5 128 GB Local development for who? How many of y'all are rocking 128GB of memory? Am I reading Apple's site correctly that it's a $10,000 laptop?
- wpm 3mo agoIt wasn't $10k a month ago
- kllrnohj 3mo agoYou don't need nearly that much RAM to run Qwen 3.6 27B, though. qwen3.6:27b-q4_K_M is only 17GB, for example.
- DanHulton 3mo agoThis is what I run on an M5 MacBook Air 32GB. Works great. I’m not having it build whole features from scratch, though. I give it pretty explicit instructions closer to the class or function level, and it still saves me an immense amount of time, while I’m very connected to the code that’s written. Definitely the sweet spot for me.
- spike021 3mo agoCertainly won't work on my M4 Pro with 24GB lol
- whynotmaybe 3mo agoI feel you! Sent from my 8gb M2 Mac mini.
- kevinrineer 3mo agoI'm still rocking my nvidia 2060, which I had purchased for $400 at the time. I struggle to imagine purchasing multiple 1k+ cards on my own dime.
- MatthiasPortzel 3mo ago
- onion2k 3mo agoNone of the examples reflect 'real work', at least not what I'd consider real work. Being able to nail a zero-shot greenfield project is relatively easy even for a small model. There's not much context to build up and it can fall back to similar examples in the training data easily. So long as you're not asking it to invent something wholly new it'll probably manage. The real test is whether or not it can work with your existing codebases. In my limited experiments Qwen 3.5 (maybe 3.6 is loads better) does OK on a Rust+React app, and less well on a C# monolith. Not to the point of being unusable but definitely poorly enough that I went back to Claude after 20 minutes. If I lost access to a cloud model and had to use Qwen instead I'd be visibly sad.
- h4ny 3mo ago> In my limited experiments Qwen 3.5 (maybe 3.6 is loads better) 1. Maybe you should tell us what those limited experiments are. 2. Maybe you should actually try 3.6 because it's huge difference in most cases. Don't forget to tell us quants and don't forget to tell us scope. 3. Maybe actually show us data compared to frontier models instead of this... vibe comment. Pretty tired of this kind of comments on HN that doesn't require logic or evidence. Just vibes. Like the pelican riding a bicycle crap that everyone has taken for granted but has no objective way of assessing goodness.
- snapcaster 3mo agoNobody owes you a scientifically rigorous write up
- sosodev 3mo agoIn my experience, even with basic project concepts the small models struggle to spin up greenfield stuff. There's just too many decisions to be made and they're not good at that. Modifying existing code is way easier if you don't expect it to be smart about it. Don't say "add X feature" and let it explore the codebase and build its own understanding. Point it at the relevant files and say "the goal is to add X feature to this code, follow Y guidelines". Now you've done the hardest part of making the decisions and it just has to follow instructions while coloring within the lines.
- anonym29 3mo agoStrix Halo user here. While Qwen 3.6 27B exhibits remarkable intelligence density, I will still take unsloth's dynamic IQ2_XXS of Minimax M2.7 over Q8_0 Qwen 3.6 27B any day of the week, and this isn't just because of generation speed either. I wrote my own custom harness, and I get hallucinated tool call parameters and bizarre invocations with Q3.6 27B even at Q8_0, but no issues with the IQ2_XXS of M2.7.
- BoredomIsFun 3mo ago> I get hallucinated tool call parameters and bizarre invocations tweaking sampler might help
- mikert89 3mo agonone of these local models are good for development, complete waste of time. nobody has $100k+ hardware sitting around at home to actually run a good model
- RedCinnabar 3mo agoCall me back when you can run these models on 16GB of RAM and any recent i5/i7. Until then, there’s no point on using these toy models.
- giancarlostoro 3mo agoYou need it to run in about 8 GB so you have extra space for the context window.
- Catloafdev 3mo agoHello, it's the internet calling, today is that day. https://github.com/ikawrakow/ik_llama.cpp https://github.com/ikawrakow/ik_llama.cpp Edit: it's gonna be slow if you're not using any VRAM. But it's possible. Software isn't going to speed that up anytime soon, it's just a hardware bandwidth limit.
- guax 3mo agoIts so funny, these "toy models" would be the wet dreams of researchers not 5 years ago. Progress marches without mercy.
- kgeist 3mo agoYeah people don't realize these "toy models" now completely destroy gpt-4o on most tasks, and no one called gpt-4o a toy model back in the day... It was OpenAI's flagship model from 2024 to 2025.
- Gigachad 3mo agoTbh in 2024 most were calling these models useless for programming and a scam. It wasn't until this year things really changed. My experience with Qwen 3.6 is it can do things, and it's super impressive it can do things, but it's not any more productive than doing it myself.
- jboss10 3mo agoThey can be ran on 32GB with 8GB VRAM. I don't think these will be on 16GB for a while. (35B MoE)
- bensyverson 3mo agoThe article is based on running Qwen 3.6 on a 128GB MacBook Pro. For reference, a 128GB MBP currently starts at $6699 USD [0] Some people will be happy to pay that premium for privacy, but at roughly 10X the cost of a MacBook Neo, that money could also buy a lot of credits on OpenRouter or frontier labs. [0]: https://www.apple.com/shop/buy-mac/macbook-pro/14-inch-space-black-standard-display-apple-m5-max-chip-18-core-cpu-40-core-gpu-128gb-memory-2tb-storage https://www.apple.com/shop/buy-mac/macbook-pro/14-inch-space...
- Insanity 3mo agoBut you have to factor in that this device will last you 5-10 years. That said, I wouldn't spend almost $7k USD on this macbook lol.
- someperson 3mo agoIn 5-10 years, incremental cloud tokens will be far cheaper (likely but not guaranteed).
- petilon 3mo agoMemory requirements of newer models will increase, so while the hardware may last 10 years it won't be able to run the latest models for 10 years.
- Insanity 3mo agoYou raise a fair point, but I'm not convinced it'll offer a meaningful difference in performance as long as we're stuck with the current AI paradigm.
- simonw 3mo agoIt can't run the latest models today - GLM-5.2 class models already need 1TB+ of RAM. ... but, the models that WILL run on 128GB (or 64GB or even 32GB) models today are a huge improvement on the best models that would run in the same amount of memory six months ago.
- seemaze 3mo agoI was interested to see that Qwen3.5-122B-A10B narrowly beat Qwen3.6-27B on Donato Capitella's SWEBench-verified-mini run with a similar 128GB UMA architecture. https://pi-local-coding-bench.dev https://pi-local-coding-bench.dev
- jononor 3mo agoMany people in LocalLLaMA Reddit community has been reporting the same, that 3.5 122B-A10B is on par or slightly better. And a 3.6 or 3.7 od the 122B is one of the models people want to see the most.
- beastman82 3mo agoFWIW I'm running gemma4 31b on my 5090 and it's pretty great as well. QAT, MTP, 128k context. I liked Qwen 3.6 27b too, it just seems that Gemma4 is a bit underrated.
- accrual 3mo agoNice. I flip flop between Qwen 3.5 9B Q6_M and Gemma4 12B Q4_K_M on a 4080 Super. They run at about the same speed and I can have them review each other's plan or diffs. For smaller projects I find them very capable, and I can step up to a better quant for slightly more challenging work.
- nok22kon 3mo agoyou can probably run Gemma4 26B on your card also at 4 bit. World of a difference compared with 12B.
- zingar 3mo agoWhere does “big model highly quantized” start getting worse than “smaller model less quantized”? Is there a general formula or is it just trial and error?
- nok22kon 3mo agopaper is a bit old, but matches current empirical recommandation: a good starting point is the biggest model you can fit at 4 bit https://arxiv.org/abs/2212.09720 https://arxiv.org/abs/2212.09720
- boppo1 3mo agoHave you tried qwen 27b q4_K_XL? It's a little bigger than the 4080 but not too much
- kofu 3mo agoMy experience also aligns with this. I'm running gemma4 31B on a 4090 through llm.cpp with unsloth models. I also run Qwen 3.6. Qwen is good for thinking and planning as it is faster, but Gemma4's generated code is much higher quality in the first try (Rust, C++ and C#). so it needs less revisions to be at a level I'm comfortable for merging.
- mbgerring 3mo agoSomething I find really confusing from this post is the MLX versions of the model running much slower. As I understand it, these model versions are meant to take advantage of Apple Silicon and MacOS APIs, and should produce better/faster results. Any insight into what’s happening here?
- verdverm 3mo agoQwen's new AgentWorld model is good too: https://huggingface.co/Qwen/Qwen-AgentWorld-35B-A3B https://huggingface.co/Qwen/Qwen-AgentWorld-35B-A3B I'm running the NVFP4 alongside Gemma4 at the same quant on an OEM Spark
- colinsane 3mo agoAgentWorld is _fantastic_. i just migrated "down" from the 122B A10B qwen model to agentworld (35B A3B) because it feels as capable, easier to steer, and it's 3x faster. also i like that if i drop more sophisticated tools into my harness (e.g. any of the NLP/RAG-based search tools in place of grep/rg), the agent will actually reach for them and make progress faster; previous models have been reluctant to embrace new tools.
- cat_plus_plus 3mo agoGemma4 31B with MTP enabled is faster and I feel a bit stronger at coding. Either one can run in 32GB VRAM or unified RAM with some tuning (3 bit weights, 8 bit kv cache)
- blopker 3mo agoI've been working with local models for the past year. There's so many possibilities, but I don't think coding is one. Coding requires so many layers beyond inference; I spent so much time trying to replicate what Claude Code does end to end locally. Understanding all the layers and keeping up with the advancements feels like a slog. Even this article messes up and misunderstands what some of the settings are doing. Qwen in particular seems to work at first, then often gets stuck in thought loops when used for actual work. However, text-to-speech, speech-to-text, and non-code LLM use cases are so useful to have local, and don't require big hardware. Having a universal reliable inference engine interface, I think, is the big unlock that needs to happen before app devs can ship these features. Personal concrete use case: meeting recording app. This uses Parakeet + Qwen to create local transcriptions and post-cleanup, respectively. Right now this app has to download and manage all these models, then bundle an inference engine to run them. It's a lot of code that probably should belong to the OS, or at least a standard interface. While apps can offload some of this to llama.cpp or a similar process over http, that's another set of setup for the user to do before they can have a useful app. Anyway, if you're getting started on a Mac, I'd suggest trying out oMLX (https://github.com/jundot/omlx https://github.com/jundot/omlx) before messing with llama.cpp. In particular they have community benchmarks so you can see what kind of performance you're likely to get: https://omlx.ai/benchmarks https://omlx.ai/benchmarks. I wished each one had more configuration details though.
- iwontberude 3mo ago> I don't think coding is one Certainly this is falsifiable easily by any of us doing it on a regular basis > Qwen stuck in thought loops This does happen when context is not managed effectively; creating plans, using subagents and compactions strategically resolves this
- blopker 3mo agoSure, local coding is clearly _possible_, but it's not practical for most people. I've yet to see a reliable setup, if you have one, I'd love to see. > creating plans, using subagents and compactions Yes, these are all things that Claude Code does for you. However, for the thought loop issue, these are not the fixes. The canonical fix is to limit the number of thought tokens (llama.cpp's `--reasoning-budget`) or try to mess with the various penalty parameters. In any case, it's not a solved problem as far as I can tell.
- doodlesdev 3mo agoI feel like I'm going insane seeing people buy these 128gb MBP for thousands of dollars to run models that are objectively much worse than SOTA and spending so much more. The amount spent on a 128gb M5 MAX can buy you a damned new car here. What the hell am I missing? Are developers in other countries living in such different worlds? (I'm aware the price is, in absolute terms, more expensive where I live compared to the USA. That reinforces what I think, because anyone sane that would've bought one of those in another country would sell them as soon as they landed here and save that money.)
- adamors 3mo agoYes they are, 6k is peanuts to a lot of people.
- JeremyNT 3mo agoI also don't understand why people in this price bracket are buying Mac laptops instead of desktop computers with GPUs? Just to flex that it's portable?
- jeroenhd 3mo agoA mac with a boatload of RAM can run models that will exceed the limits of any GPU not worth at least twice the Apple hardware itself. You get fewer tokens per second, but at some point the balance between quality and quantity makes the large model size worth the spend. When you're spending this kind of money, you may as well treat yourself to a pretty screen and some decent speakers. Nothing the competition doesn't offer these days, but you get them for free with the car-priced RAM upgrade so why go for less.
- ilogik 3mo agoWhat GPU can I buy with >100GB of memory?
- verdverm 3mo agoDGX Spark is one, but really depends on how much you want to spend
- jjcm 3mo agoI'd also look at the qwopus distil if you're using qwen 3.6 27b. It's a nice refinement of the current 27b with slightly better stats. Jackrong has a few different ones available depending on what you're trying to do: https://huggingface.co/Jackrong https://huggingface.co/Jackrong
- markdog12 3mo agoI've tested it extensively for actual local development for my job, and hard disagree here. It's a waste of time to use it. Wish it were not true.
- beastman82 3mo agoI posted elsewhere but if you have more space try gemma4 31b
- CurbStomper 3mo ago[dead]
- SkitterKherpi 3mo ago27-30B in general seems to be the level where you actually start having decent models. I just wish consumer hardware hadn't stagnated so much that we can't easily go higher than that, and that even running those requires a $5k machine now.
- dmezzetti 3mo agoLocal models are great for a lot of things past just software development. We need to move towards solving other real world problems vs just building software. I've been focused on that with TxtAI (https://github.com/neuml/txtai https://github.com/neuml/txtai) for 6 years now.
- Otternonsenz 3mo agoIs there any hope for people that cant even run 27B parameters, Qwen3.6 or otherwise? Are there any quantized models that do well with tool calling at smaller parameter sizes? I do not have a crazy rig, a modest gaming one at that, but in trying to understand more about agents and their capabilities, I am SOL with my 16 GB of RAM and 8GB of VRAM. I can get most small, non tool calling models to perform well, but I've had major issues with anything over 9B doing anything more than reasoning (egregiously slow at higher parameter counts). And so far, I cant get even Pi to extend itself or do any meaningful work with any of the models I currently can get to run.
- jadbox 3mo ago[dead]
- fumeux_fume 3mo agoI suspect with those specs, you're not in the game right now for reliably using local models for code generation. The easiest way in is a MacBook with at least 32GB of RAM. This should be able to run a 4bit quantization of qwen 3.6 using the MLX format really well.
- Otternonsenz 3mo agoNow that I’m dipping more into this space, am gonna see what I can upgrade with the motherboard I have, but RAM pricing as it is, I’ll need to be smart about when I upgrade. I very much appreciate the frank response, as it makes me feel less defeated at knowing my understanding of how it should work is not the full issue, hahaha
- fumeux_fume 3mo agoM series macs are usually used for running these LLMs locally because the GPU and CPU share the same pool of RAM at very low latency. If you upgrade your RAM on a different kind of chipset without the Unified Memory Architecture, then it'll be much slower to produce all the tokens you need. Just another data point to add to your upgrade equation.
- suthakamal 3mo ago[flagged]
- prasanthabr 3mo agoHas anyone considered a home server? Assuming mobility is not important if we pick components to match a similar hardware would it be more value for money?
- LeBit 3mo agoWhich components are you thinking about?
- prasanthabr 3mo agoAm unsure - was hoping someone tried this and there is a tested component list of consumer grade pc parts that can do the trick
- drillsteps5 3mo agoA decent gaming machine perfectly doubles as your friendly local inference server. Just start llama-server with the model of your choosing and start chatting with it through its Web interface or connect any chat completion-compatible client (agentic or not) which will use REST to send requests and receive responses. From any device on your network. Voila.
- cpburns2009 3mo agoGenerally speaking a home server/workstation set up is going to provide better performance at lower cost. You don't sacrifice much mobility either so long as you have an internet connection and can either SSH tunnel or use Tailscale (never used, just know it's popular).
- Greenpants 3mo agoI specifically chose a Mac Studio 128GB as my home server that's also running LLMs to be always online, in part due to the minimal idle power consumption and mostly fan-less operation. It's definitely expensive, especially nowadays, but I can still recommend Mac Minis as a cheaper alternative for someone to just get started with an affordable, always-on home server that won't annoy any housemates. I think both are in some sweet spot in terms of value for money, depending on what you're looking for in a home server. If image or video generation is your thing, look further though, definitely look into a proper GPU then. Macs are quite slow at that. They're just great at MoE LLMs because it's mostly a matter of (V)RAM size.
- dom96 3mo agoWhat do folks use to keep on top of new model releases that are appropriate to their system? i.e. the models that will actually work on the MacBook Pro with 48GB of RAM or whatever their specs are. I've seen sites here and there but they feel like quick little toys that don't get updated, so they always suggest old models.
- IronWolve 3mo agoI think things are moving fast, tested that new vibethink-3B, works on many small tasks/fast, and playing with ornith-35B with a draft vibethinker-3b as a draft gave me some good speed/results. Was just trying to see how small I could go and get acceptable results, but yeah, larger Qwen 3.6 with MTP is going to be better. Cant wait to see how AI model (unsloth/local-llm/heretic/reaper/etc communities) are tweaking/engineering quality down into smaller models. Lots of new things coming out.
- blueside 3mo agoi have been trying several open source models for the last few years. running qwen 3.6 27b on my 4090 is the first local llm i have used that made me start to second question if anthropic and openai are actually worth the (already) insane valuations. don't get me wrong, the frontier models are leaps and bounds ahead of what qwen/kimikgemma are doing - but i don't need to drive a ferrari to the grocery store everytime either.
- mannyv 3mo agoFYI token speed is somewhat irrelevant for agentic development. You let it run, then you come back. The whole point is that it's asynchronous. If it takes 4 hours, 8 hours, 16 hours...who cares?
- iagooar 3mo agoI love my MacBook Pro M5 128GB RAM and I love qwen3.6. BUT DO NOT buy this MacBook if you plan on doing serious coding using local LLMs with it. The reason is simple: your fingers will burn and your head will explode from the noise. Running any kind of sophisticated job on the very laptop you are using is just not viable. Sure you can use it in clamshell mode, but forget touching it while working with AI coding or agents. If you want to run Qwen3.6 27B / 35B at its best, get a MacMini M4 with 64GB of RAM and put it in the basement - or at least a few meters from your desk. Connect to it over LAN or Tailscale. The MacMini will also cost you almost 1/3 of the MacBook Pro. Thank me later.
- zkmon 3mo agoThe Q6_K gguf fits nicely on a 24GB GPU. That's amazing.
- singpolyma3 3mo agoWith 128 you can run 122b ;)
- verdverm 3mo agoGet an OEM Spark instead, mine are silent and can fit 2 qwen/gemma at 8bit or give you room for a bunch of other, smaller models (embed,rerank,etc)
- oceanplexian 3mo agoIf you want to do coding with a local LLM your best bet is a 6 year old Nvidia 3090 which is substantially more powerful than the highest end overhyped Apple product for 1/5th the price.
- chorizo 3mo agoThat’s 24GB VRAM. Not enough to run a 27B model at a useful quant+context size.
- SkitterKherpi 3mo ago
- narrator 3mo agoIn hindsight, the Mac 512gb for about $10k was a total steal given that to run GLM 5.2 you need a 4x H100 to get the necessary amount of VRAM. Yeah the h100 is 2 to 8 times faster, but it's $20k a month to rent a 4xH100 VPS.
- MangoCoffee 3mo agoRunning LLMs locally for development doesn’t make sense to me. The hardware gets outdated in just a few years. Even hyperscalers replace their GPUs faster than they can buy them, plus the cost of running it locally, isn’t cheap. the cost saving just ain't there.
- guax 3mo ago> replace their GPUs faster than they can buy them How does that work? They have negative GPUs now!
- logankeenan 3mo ago3090 was released six years ago and is still very relevant for running models locally.
- jboss10 3mo agoQwen 3.6 35B runs on 32GB with a 1080. That GPU is from 2017.
- kgeist 3mo agoFrom the perspective of LLM inference, you currently mostly care about: - Memory bandwidth; BUT the requirements are currently capped because models have stopped growing at around 1-1.5 trillion parameters for quite a while now. You only need more bandwidth if you're optimizing for the highest possible concurrency (i.e. you're a cloud provider). Also, MoE exists. - Support for native low-precision math (like FP4 and FP8); BUT once your GPU supports native FP4 (Blackwell+), there's generally no reason for GPUs to go lower because of the obvious quality degradation. - VRAM capacity - just like memory bandwidth, it's practically capped by 1-1.5 trillion parameter models and is unlikely to need much more in the near future. Also, the current trend is toward miniaturization: modern 30B-class models (which require far less VRAM), now completely destroy 200B-class models from just two years ago on most tasks. We also have better understanding now how to compress contexts. Most model improvements currently seem to come from RL/harness-based methods, not from scaling models or running new algorithms that require fundamentally new GPUs. So I don't see why GPUs that exist today must become "outdated" in a few years. They'll be seen as outdated by hyperscalers because they need to serve the maximum number of users as cheaply as possible, so of course they'll replace their GPUs with newer ones that have higher memory bandwidth or more tensor cores. But you don't need that for local inference.
- zx76 3mo agoI see a lot of people writing about how expensive the hardware to run these local models is - but see no mentions of the Intel Arc Pro B50/B60/B70 which seem like decent value if you're not interested in Apple kit (as much as anything can be decent value in the current status quo). I just got a B70 with 32GB RAM for the equivalent of $1200 (incl. sales tax and import duties to my non-US location, so presumably it could be cheaper elsewhere). The memory bandwidth is 608 GB/s. For M5 Max (32-core GPU) it's 460 GB/s and for M5 Max (40-core GPU) it's 614 GB/s. A 3090 is still faster at ~900 GB/s but you're getting 32GB VRAM for a lot less than equivalent Nvidia cards. It's about 1/3 the bandwidth of a 5090 for 1/3 the cost, but with the same 32GB VRAM. If you're interested in being able to run bigger quants with some context and stay on a lower budget then it's an appealing trade off. I'm still exploring using these local models so don't want to spend the equivalent of $5 000 - $10 000 just to test it out. I don't mind slightly slower perf to do some experimentation more affordably. I actually got an B50 16GB (with meager 70w TDP!) first to test an Intel card with my stack - it worked easily with Ubuntu & Vulkan. I'd read a lot about hassles and people writing them off as unusable but it seems like these are often with SYCL which doesn't even seem to outperform vulkan and so why bother? (The B50 was just $370 inclusive tax and duties). Literally `apt install` the vulkan libraries and it worked with default xe driver in 26.04 and the vulkan build of llama.cpp. The SR-IOV PF/VF also just works with qemu/kvm, no tricks required. Since I got it fwupdmgr has updated the firmware twice so Intel is presumably actually trying to support these products.
- kristianp 3mo agoInteresting that Intels latest consumer GPUs only have 10 and 12GB respectively for the B570 and B580.
- bblb 3mo agoI got B70 few days ago. Running on CachyOS. 9070XT on PCIe x16 and B70 on the x4. ROCm nightly was pretty easy to setup and get up running. The 9070XT has been a decent card for my use cases. But the SYCL ecosystem versions. Absolutely horrendous and everything is hundred commits behind. Vulkan is probably the only way forward with this card.
- 3mo ago
- diseasedyak 3mo agoI have 24GB of VRAM (via a RTX 4090) and run Qwen3.6-35b:iq4, so it's importance-aware quantization and isn't nearly as dumb as it sounds like, fitting the 35b into 18 GB so you have some left over. So far I've had no issues, other than it taking a while for things like image gen, which I found out if you're gonna do with any alacrity, just have a cloud model do it. For anything else local, including writing some automation scripts and such, it works great.
- ai_fry_ur_brain 3mo agoWhats your example of a "great automation script"?
- Zambyte 3mo agoCan you link the model? I also have a 24gb card (7900 XTX). I've been having success with the dense 27b model, but I'd like to see if the 35b iq4 is any better.
- jboss10 3mo agohttps://unsloth.ai/docs/models/qwen3.6 https://unsloth.ai/docs/models/qwen3.6 And https://huggingface.co/collections/unsloth/qwen36 https://huggingface.co/collections/unsloth/qwen36
- drillsteps5 3mo agoI honestly don't get the hostility against local models in this thread (and in some other threads recently). I haven't seen anyone make an argument they are as good as SotA (OpenAI, Anthropic). It's just they are approaching state where they are "as good" for some _limited_ set of use cases. Which will allow us to resolve 2 primary issues with these SotA models: privacy and vendor lock-in. Plus, they're very useful for education purposes, you get to explore what things looks like under the hood, play with various models, tools, maybe put something simple together yourself. You get Macbook - great. You got gaming rig with a decent GPU - great (set it up as a dedicated server that you connect to through simple REST). What exactly is wrong with any of that?
- simplyluke 3mo ago> I honestly don't get the hostility against local models in this thread Consider that there are literally trillions of dollars being wagered on this not being the future state of computing. Not even speculating that HN is being astroturfed (though I see no reason it wouldn't be by interested parties), but many of the US tech employees here have direct financial incentives in various forms to be rooting for the failure of open source and optionally local models.
- starefossen 3mo agoWe have have had the same experience (qwen3.6 rocks) when we are evaluating local models for our developers in the Norwegian Government https://github.com/navikt/mlx-workspace https://github.com/navikt/mlx-workspace
- cdnsteve 3mo agoCheckout details on what this runs on for local AI here: https://tokenstead.ai/models/qwen3-6-27b https://tokenstead.ai/models/qwen3-6-27b
- ljosifov 3mo agoRunning 27B dense model on M5 128GB is ok, but one can do better. On M5 128GB one can make use of the ram and use sparse MoE. For example, DeepSeek-V4-Flash will fit, served by DwarfStar (https://github.com/antirez/ds4 https://github.com/antirez/ds4). One will probably improve 2x the token/sec speed, given DS4F 13B activated params in the MoE are ~1/2 of the ~27B of the dense Qwen. 27B Of the Qwen fit even on a cheaper 24GB card, e.g. amd 7900xtx (<$1K?) or slightly dearer nvidia 3090 (with cuda). With ~900 GB/s bandwidth they will likely be ~50% faster than the M5 with 600 GB/s.
- drnick1 3mo agoWorks beautifully on a 3090, very usable speed. Don't expect Opus 4.8-level performance, but there are some things you just need to keep local.
- ljosifov 3mo agoTrue - they are workhorses. Not super bright, but good enough for lots of everyday tasks. I've found sweet spot to be turning thinking off, as it adds small or no value, while increasing the token count and waiting time. Last 27B I used was https://huggingface.co/Jackrong/Qwopus3.6-27B-Coder-GGUF https://huggingface.co/Jackrong/Qwopus3.6-27B-Coder-GGUF - specifically post-train adapted a bit to run with thinking off. I saw today the 35B-A3B MoE from the same HF acc is out, downloading that rn to try.
- ShizuhaLabs 3mo ago[flagged]
- jboss10 3mo agoI don't understand the talk about how expensive the hardware is. These models can run on very old or old and low end. I've been running Qwen3.6-35B Q4 on an old 1080 GPU(8GB vram) with 32GB sys RAM. I have a i7-12700. It does about 30 tok/s which is enough for me. It's about half what the online models do, but it's enough. I've heard their 9B models are also good, but they aren't much faster if you have the ram and a nice cpu. These qwen3.6 models are the first ones I find can do much. GPT OSS was good, and Gemma4 is better. Gemma knows more facts, but qwen3.6 is smarter.
- felooboolooomba 3mo agoMind sharing the command line you use to rig it up?
- CMay 3mo agoThe MoE models hold up better on old hardware, but the dense models like this post promotes are in fact better. This isn't unique to Qwen. Are the dense models better-enough to use given the performance costs? It depends on what you are doing. If a model runs fast enough for your use case and does exactly what you need it to, then you don't need a much slower model that might be more accurate. If you do anything more complicated, the dense models become more necessary and they are much more computationally heavy by comparison. On your hardware an Unsloth quant of Gemma 4 26BA4B QAT would likely give you better results, but because it has 4B active parameters instead of Qwen's 3B active parameters, it will probably run slower.
- jboss10 3mo agoI should try gemma4 more for coding, since qwen3.6 and gemma4 came out I've focused on qwen. For earlier releases I found qwen was smarter, but gemma had more knowledge. But for coding I always want it to learn how to do the task, not just assume/halucinate.
- felooboolooomba 3mo agoWhat's the minimum requirement for a Nvidia card to run it? For let's say 10 t/s.
- ctkhn 3mo agoI have been running qwen 3.6 35b a3b with opencode on my macbook pro 16" with m3 max and 64gb ram, and it's been great for local planning and coding. To be honest I have been on and off wishing I had future proofed with the 128gb after seeing how powerful 64gb is. On the other hand, I also haven't run up against a wall with a model that is just slightly larger than qwen.
- Xeoncross 3mo agoWhat is the speed on responses? (t/s) The full 128GB is surely helpful in keeping browsers, editors and other things running since even 20-35GB models + k/v caches can eat up a lot of the core 64GB in my experience.
- LeifCarrotson 3mo agoI've also been running Qwen 3.6 35B A3b on my Windows laptop (64 GB RAM, a 4GB GPU) and it's at least tolerable. It's not fast - a few tokens per second, slower than reading speed - but I can give it a task and come back later. That was a $600 laptop off eBay a few years ago, not a $6,000 machine. Are these unified memory Macs and giant 24GB desktop GPUs achieving dozens or hundreds of tokens per second commensurate with their 10x-20x cost?
- jaggederest 3mo ago35b A3b runs ~100 tokens a second on the best M5 Max gpu setup.
- ctkhn 3mo agoI got around 50-60 on my m3 max so 100tps seems very realistic for 2 gens later of chip and double the ram
- Getchowned 3mo ago[dead]
- cpburns2009 3mo agoBefore you run and go purchase a unified memory computer (e.g., DGX Spark, Mac, Ryzen AI Max 395 / Strix Halo), be aware dense models generally run slow on these machines. Dedicated GPUs run dense models significantly better. Look for benchmarks for your prospective machine. If you really want one of these, you'll be better off running Qwen 3.6 35B or another sparse MoE model.
- alansaber 3mo agoIs qwen finetuned/RL'd on any agent harness? Or does it just work well enough off the bat with opencode?
- cpburns2009 3mo agoIf Qwen is finetuned for a hardness, it'll be Qwen Code. Qwen 27b works well enough in OpenCode though which is what I use. My one complaint is it likes to get cute with bash commands instead of OpenCode's built-in tools. I use a skill to steer that.
- sometimelurker 3mo agoworks well on pi, but I use a smaller one at higher tok/s for more repetitive tasks
- devin 3mo agoIf I have 10k to spend, what should I buy for the best local model experience?
- simplyluke 3mo agoI really think giving it a year for the hardware market to come back to earth and spending a fraction of that for API access to the same models is a better use of the money.
- devin 3mo agoImplicit in your answer is the belief that they will come back to earth. I wonder how realistic that belief is.
- simplyluke 3mo agoWe have decades upon decades of hardware getting dramatically cheaper year over year for the same performance, and ~1 year of the inverse due to dramatic buildout for AI. It's a surprising example of the recency bias to me to assume anything other than the market returning to its historic norm, even if the AI buildout doesn't slow, producers will scale factories to meet that demand.
- devin 3mo agoI look forward to re-evaluating this statement in, what do you say, 12 months from now?
- simplyluke 3mo agoI’ll toss $10k in the s&p and you buy the rig and we’ll see who feels like they made a better call?
- devin 3mo ago10k in the S&P is by default a far better investment than some computer components. You could say the same thing like "you buy a car and I'll put my money in the S&P and we'll see who's happier in N months". We were speculating on the cost of components going forward, not on whether the S&P is better place to park 10k.
- zedascouves 3mo agoJust tried on some arduino code. after 10 minutes i got a list of improvements to my code. I ran those throu opus saking if it was good advice and was not impressed: I read the actual qr_scanner.ino. Short answer: partially, but I'd push back on most of it. That review reads like generic ESP boilerplate advice written against an imagined version of your code — several of its "fixes" are already in your file, and its headline "critical" claim misreads what the code does. Going point by point:...
- pkroll 3mo agoSince no one else posted it... I have open-webui pointed at a linux box with 128 gig of ram and an RTX Pro 6000, and after a couple of runs on trivia, had it do one of Open WebUI's conversation starters: "Show me a code snippet of a website's sticky header in CSS and JavaScript." 72.06 t/s. That's the full Qwen 3.6 27B model BF16, using MTP, running on Ollama. Yes I know I should bite the bullet and get vllm running on that box. That was, also, at a 570 watt limit: I normally run a little less, but when I first tried this I actually forgot I had set the limit to 300 (it's a hot day, I figured why fight the A/C?), and at 300 watts the same question came back at 69.38 t/s. (The extra power matters more for compute bound things, the difference in generating LTX2.3 videos is considerably higher... but still not linear.)
- marcuskaz 3mo agoWhen is Amazon Bedrock going to get these newer models? Offloading compute to them is much easier, except its still a limited set of open models. Most companies are already running in AWS, so it's an easy add, models run in a trusted location, just another line item on the Amazon bill. You don't have to talk anyone into signing up with a new vendor. Plus you don't have to worry about local hardware at all.
- zerolines 3mo agoYup, been rocking theQwen3.6-35B-A3B-MTP-GGUF locally with 88tk/s it's amazing.
- deleted 3mo ago[deleted]
- mips_avatar 3mo agoI think the sweet spot right now is 2x 3090s and a pcie 4 motherboard with 64-128 gb of ddr4 ram, you can build this right now for $3k and it runs qwen 27b/35b stupid fast at int4.
- tasoeur 3mo agoI know how to build PCs but suck at picking parts, would you happen to have a recommended build or links to people who've done similar ones? Heck I'll click on an affiliate link to support the author of the build :-)
- mips_avatar 3mo agoHere's my build! https://jonready.com/blog/posts/local-llm-rig.html https://jonready.com/blog/posts/local-llm-rig.html I love it because the watercooled 3090s are completely silent even under load. Facebook marketplace is definitely the move for a lot of the parts unfortunately, since you ideally would have higher end parts that are 2-3 years old.
- christoff12 3mo agoI just burned 20 minutes because I wanted to play hex minesweeper: https://hexabomb.pgpln.app https://hexabomb.pgpln.app Source: https://chatgpt.com/share/6a42dd8a-4e28-83e8-9ef7-6ba56d665c95 https://chatgpt.com/share/6a42dd8a-4e28-83e8-9ef7-6ba56d665c...
- stared 3mo agoNice! If you want to play a hyperbolic minesweeper, Hyperrogue features that https://hyperrogue.fandom.com/wiki/Minefield https://hyperrogue.fandom.com/wiki/Minefield
- v3ss0n 3mo ago3.5 122B is much better. 27 B is bad at Long context and Svelte
- hollowturtle 3mo ago> Real work Ok that's the part I'm interested in, don't care about minesweeper clones.... > Make a landing page selling candles for women that are into wellbeing and SPA. can't be serious...
- recursivedoubts 3mo agoI would like to offer someone the next openclaw: a GUI for the mac that allows people to manage and install local models with a single click, provides GUI tools for tweaking important aspects of them, and also provides a good command line interface to those models.
- hollowturtle 3mo agoollama is a good starting point
- simplyluke 3mo agoThe open source models have gotten heavily conflated with local development. While that is cool and I'm excited about the future of local LLMs, it is not necessary to play around with these models. Without shilling for companies I don't have a relationship with, there are a number of companies who will give you an API just like Anthropic/OpenAI and you pay per token, albeit much cheaper than the frontier labs. I've been using the full GLM 5.2 model this way (through opencode) at work for the past week. It's quite impressive.
- Alifatisk 3mo agoShouldn’t we call them open weight models?
- simplyluke 3mo agoThat's probably more precise.
- mark_l_watson 3mo agoI can come close to agreeing because queen-3.6-27b is my second favorite for local coding. I am using gemma4:26b-a4b-it-qat-48k (the "-48k" is from my modifying a model run with Ollama to always use a 48K context size). On a 32G Mac I use gemma4:26b-a4b-it-qat-48k and OpenCode and on my 16G MacBook Air I use gemma4:12b-it-qat-16k ("-16k" is my resizing context size) and little-coder. I break up projects into small libraries because local coding works better for me using small code bases. I find that for local coding, I need to spend a lot of time building concise SKILLs for specific things I work on and try to only enable one or two skills per coding session. To the author of the linked article nice job, and if you feel like adding to it, please add details on your setup.
- brandall10 3mo agoCurious why OpenCode instead of a more 'full-fat' version of Pi with the larger model? I feel like the amount of context bloat that OpenCode puts these small models into the dumb zone too quickly. The system prompt alone is 9k tokens, and when you add your own setup it can easily creep up to 15k.
- mark_l_watson 3mo agoI disabled many built in skills and increased the context size. I also use little-coder that is based in pi.
- hoppp 3mo agoIts feasible but that laptop is not available for most devs. I do have access for a 64 gb ram mac mini but most people don't.
- XCSme 3mo agoConsidering the cloud version, all three models compared in the article (Qwen 3.6 35BA3b, 3.6 27B and DeepSeek V4 Flash), have very similar performance[0], BUT on cloud, for some reason DeepSeek V4 Flash is 10-20x cheaper than the Qwen models. If Qwen models are so much easier to run, why are the providers charging more than V4 Flash? [0]: https://aibenchy.com/compare/qwen-qwen3-6-35b-a3b-medium/qwen-qwen3-5-27b-medium/deepseek-deepseek-v4-flash-high/ https://aibenchy.com/compare/qwen-qwen3-6-35b-a3b-medium/qwe... <-- compare how the three models draw hamsters svgs, lol
- Gigachad 3mo agoAlso confused by this. Deepseek V4 flash is so much better than Qwen 3.6 yet cheaper to use.
- jingpostmedia 3mo ago[flagged]
- jboss10 3mo agoLook into deepseek's papers. They have done some stuff recently about improving inference and it seems to be how they can sell tokens so cheap.
- SamInTheShell 3mo agoThis is probably the first small model I got through some simple web game tests without having to reset the context. It tends to opt to overwrite an entire file instead of doing edits... which editing is where most of these small models fall apart along with getting stuck in repeating loops. Only 24k tokens in so far, it did some decent newbie work.
- cloudengineer94 3mo agoI'm using Qwen and Gemma 4 locally and it's pretty good stuff, not frontier level but gets the job done.
- sourcegrift 3mo ago[dead]
- max8539 3mo agoRunning this model on a 48 GB memory MacBook Pro when offline, it performs its tasks, but of course, it’s slower than Claude or Codex.
- mashygpig 3mo agoIt's fun to run a model locally, but I don't think the economics make sense for anyone just trying to use models atm. It's absurdly cheap to use the same model via openrouter in comparison. Seriously, just put $10 into openrouter and play with models that are cheap but bigger than what you'd reasonably be able to run locally like deepseek v4 flash (unquantized). You'll be surprised by how far that $10 goes for a model better than what you'd be able to run. Even further on the model you would be able to run locally. Then think of how many long it would take to match the cost of spend + power on doing it locally...
- SchemaLoad 3mo agoAgreed, I'm waiting for the time when 48GB+ ram is just the standard that computers come with rather than being the absolute top tier option. It just doesn't make sense to spend extra on a local AI computer right now when the same money would last for a decade of API pricing.
- boppo1 3mo agoHave you considered this may never happen? What if datacenters continue to swallow all capacity?
- Saris 3mo agoEven with deepseek v4 flash I burned though $5 in credits in a day just playing around with Hermes, and qwen 3.6 35B is significantly more expensive. I can run qwen 3.6 35B on my gaming PC at around 50 tok/s and other than power cost of a tiny bit extra per month, it's hardware I already owned from years ago. I'm not really sure why qwen 3.6 35B is so expensive on openrouter, it seems abnormally high for what hardware it takes to run it.
- myzek 3mo agoHow do you run 35B on a gaming PC? I'm trying to go the same route, but I have a 5070Ti with only 16GB VRAM (I bought it for gaming) and I'm not sure how to run anything decent on it. I have 64 GB RAM if that matters
- blagui 3mo agoHow you can do dev in 2026 using 64k context and without sub agents? The benchmark seemed fine until I saw that. If you use sub agents, they will overwrite the cache and each request will trigger full reprocessing. Have fun with that as it will crash the t/s metrics on each prefill on top of the max 64k including input + output is a major blocker. If you push the context higher and add parallel slots the requirements will be far higher and the numbers less shiny.
- so_it_be 3mo ago[dead]
- deleted 3mo ago[deleted]
- dhanush_2905 3mo ago[dead]
- drnick1 3mo agoHas anyone managed to cleanly integrate Web search into local models (run with llama.cpp)? The biggest limitation of the class of models that fit into one or two consumer GPUs is that they lack world knowledge, but presumably this can be remedied by enabling access to use the Internet.
- kroaton 3mo agoYou're late to the party, mate; we've been doing this for years. Grab a SearXNG instance, stand up an MCP server for it, and expose the tool into your system prompt. Or use Brave Search. Or Exa if you want to pay. Any of them work. The model will pick it up straight away. Even llama.cpp's bundled web UI handles it fine. Dead simple.
- drnick1 3mo ago> Grab a SearXNG instance, stand up an MCP server for it Which MCP server do you use?
- mwowow 3mo agoWorking fine with LM Studio + Web search plugin
- Havoc 3mo agoSearxng is the ghetto solution. Commercial uruky is good. Basically Kagi except you can also run api calls over it Neither is going to return much knowledge. Basically just relevant url so you need a second tool to grab them and there bot walls get tricky
- macwhisperer 3mo agohi guys... I run specialized quants on my 24gb air.. (I specialize in 3-bit quants that punch above their weight).. try out my version of 3.6-27b I think you be impressed https://huggingface.co/macwhisperer/Qwen3.6-27B-SuperDense https://huggingface.co/macwhisperer/Qwen3.6-27B-SuperDense
- macwhisperer 3mo agoalso for those with only 16gb-- try this model https://huggingface.co/macwhisperer/Gemma4-12B-SuperDense https://huggingface.co/macwhisperer/Gemma4-12B-SuperDense its exceptional!
- Reuben_Santoso 3mo ago[dead]
- senorqa 3mo agoOn AMD R9700, I'm getting ~90 t/s with 35b MTP variant and ~40t/s with dense 27b MTP
- taf2 3mo agoBest way to make your M series macbook pro feel like a good old fashion intel macbook pro. Run a local model.
- kristopolous 3mo agoHelp me improve local model performance with petsitter! It basically exploits the face that time can be traded for intelligence with local models https://github.com/day50-dev/Petsitter https://github.com/day50-dev/Petsitter
- trey-jones 3mo agoQwen3.6 was the first model I ran locally that seemed smart, but qwen3-coder:30b is way, way more responsive and more accurate for writing code according to my tests, including human-eval. If you can run one than you can almost certainly run the other. If you haven't tried qwen3-coder I would definitely recommend it. It feels positively snappy compared to every other local model I've tried. All you need is 32G VRAM and some heat dissipation.
- androiddrew 3mo agoDual AMD Radeon AI Pro 9700s (600 watts total 64GB of vram) runs Qwen 3.6 27B at FP8 with mtp on vLLM at 50ish TPS for decode. Cards cost $1300 a piece. Enough KV cache to fully max out two concurrent sessions. It was super rough going to get started with them back in January, but right now the cards purrrr and I haven't even tried tuning yet. You need to use a patched vLLM image with aiter but besides that things are finally working on the ROCm front.
- ThunderSizzle 3mo agoAgreed. I have a single 9700 and I'm able to fit Q6 27B at 30tps or Q5 35B at 100tps very easily via llamacpp running vulkan. The results are impressive considering the amount of people trashing AMD and still trying to recommend 3090s. I hope to buy a 2nd one at some point, but I also hate the version hell of vLLM, the R9700, the ROCM version, and Qwen3.6 all not agreeing with each other. I haven't gotten vLLM to run properly for Qwen3.6, since the version that runs on a 9700 doesn't support 3.6 yet. I'm trying to quickly hack out a optimized path for just Qwen3.6 to run against rocm natively (e.g. my own inference server for 9700s basically) and see if it can perform better than llamacpp vulkan's results. Word of caution - the last llamacpp with good performance was b9209 from a month ago. After that, for some reason, vulkan performance dropped by 10x, which has made me lose confidence in llamacpp in the long run. Having said all that, 3x is 96GB for 4k and peak 900 watts. A 96GB Blackwell is $12k and peak 600 watss. And they will have a similar memory throughput (minor negative to the AMD cards for split processing). It's crazy how price efficient the r9700 is compared to the Nvidia cards.
- this_was_posted 3mo agoI'm getting around 45 tps on a single r9700 for Q6 27B with build b9811 ( using https://github.com/kyuz0/amd-r9700-ai-toolboxes https://github.com/kyuz0/amd-r9700-ai-toolboxes ) with the following parameters: llama-server -hf unsloth/Qwen3.6-27B-MTP-GGUF:Q6_K -c 135000 -ngl 999 -np 2 -t 16 --temp 0.0 --top-p 0.95 --top-k 20 --min-p 0.00 -b 4096 -ub 4096 --chat-template-kwargs '{"preserve_thinking": true}' -fa 1 --spec-type draft-mtp --spec-draft-n-max 2
- jimmaswell 3mo agoMy partner has been trying various models on our server but we haven't gotten anything to run at a usable speed. Q30H engineering sample (Xeon 8570) with two cpus, 56 cores per CPU, 768GB DDR5 RAM running at 5600MHz, two old 3090s in it at the moment with an NVLink and we could put our third in there. We built this server before the prices skyrocketed because we happened across some Tyan boards on Woot that were absurdly cheap for what they are (the motherboards should be $1000+ but we got them for a few hundred). This thing sounds like it should be a monster but we keep running into issues of the old GPU architecture, lack of support for AMX or AMX not being as big of a help as you'd hope when it does work, etc. Apparently we only got 5 tokens per second trying to set up Qwen 3.6 27B, and a similarly bad result trying to run GLM 5.2 which fits in memory but the custom kernels we had to try to contrive were too slow. I feel like this system should have tons of potential, especially if something was designed to let the AMX and huge system memory shine. Does anyone have any suggestions? This thing was fun to set up and it's really cool but it's been a bit disappointing not getting any big tangible results so far. We have a similar system on a single-cpu Tyan board with 256GB RAM that I'm hoping we might be able to use in conjunction with the first one if EXO ever gets good Linux support for GPU/RDMA over InfiniBand.
- christina97 3mo agoStart with a quant, you can run the Qwen 27B model at 4-bit on one 3090, presumably 6/8-bit on 2x3090.
- danielrmay 3mo agoYes, this should be a monster machine. Ampere is an older generation, so I expect that's where some of your issues have been
- fossheart 3mo ago> I recommend llama.cpp - a direct, open source tool that allows running models on various devices. You don’t need Ollama, and frankly - I would recommend against using that on ethical grounds. > https://sleepingrobots.com/dreams/stop-using-ollama/ https://sleepingrobots.com/dreams/stop-using-ollama/ I had faced roadblocks while integrating with openclaw using ollama (Was trying to experiment with `qwen3-vl:2b`). I was tracking the issue back to openclaw at that time, I didn't even consider investigating ollama. I attached a threads post here where I'm talking to meta ai to expand on both scenarios (not to use ollama, but llama.cpp & my take on the why this is the way it is - ie. commercial gains) https://www.threads.com/@riojos/post/DaMXIs4k4w8 https://www.threads.com/@riojos/post/DaMXIs4k4w8
- m3kw9 3mo agoHmm, i used it and it can't get past a simple coding test that 5.5 passes with light reasoning
- kopirgan 3mo agoLost count of number of times I read this or similar: For me it’s the first local model that actually makes sense as a general intelligence.
- Nasser_CAD 3mo ago[flagged]
- happyash1 3mo agoQwen is so good a model.
- rvz 3mo agoWhen reading the comments, it seems that in the AI race to zero, Apple was already at the finish line. as predicted. So it will be no surprise that there will be a time where everyone will be able to run a local model, say GLM 5.2 locally on their machine. Like it or not.
- fabijanbajo 3mo agoWe need machines designed around wide memory + sustained inference thermals, not gaming/creator chassis we're borrowing. Until then "local dev" means clamshell + external fans.
- deleted 3mo ago[deleted]
- decide1000 3mo agoA lot of replies here are about Mac devices and their support for these 27B models. I own a MacBook but use a Lenovo Thinkstation PGX to run my models. It has a gb10 Blackwell gpu and 128gb unified memory. You can connect multiple ones.
- cloudcanalx 3mo ago[dead]
- modgate 3mo ago[flagged]
- konart 3mo ago>Real work This part should have featured something about real work. But instead it features a paragraph about one-shot bs that creates "something". Unless your work is to create thousands wordpress tremplates to sell - this is not a "real work". Give it a repository (any kind of OSS project will do for an example) and a github issue requesting a knew feature or describing a confirmed bug. (you can and probably should write a prompt for LLM shough, don't just provide the issue itself) And then whatch it go. And then judge the result and it's quality. Sorry, but from my experience 27B is just useless. You do get a result and some times it does work, but most of the times it is not event on junior dev level. And it takes it a lot of time to do the thing, unless you have an extremely expensive machine.
- hypfer 3mo agoIf your expectation is to treat it as a coworker, then you're right. If your expectation is to treat it as a tool, then you're wrong. I guess that's where the disconnect lies.
- konart 3mo agoDefine "a tool" for me and we can talk. I already have tools for autocomplete, working with structured data and many more. Deterministic tools. Obviously you do not expect something like that from a model with some harness. It can read some input (user's or other tools) and give you some output. My expectation is that this tool, given some meaning full input (instructions, expectations, motivations and an optional source files to work with), will produce something that will actually be aligned with the input. For example: consider I have a services that has some sort of events created now and then. I what those events to be available for other services. So I decide it to have a transactional outbox and an observer that will pull events from the outbox and put them into a kafka topic. My expectation is that I can give this tool some context (source code and description), state my instructions, expectations, motivations, design decisions and have an implementation as a result. My other expectation is that given my context etc and agent's context (skills etc) were correct and adequate - the outout will also be correct and adequate.
- agenticup 3mo agoqwen 3.6 27b and qen35b a3b work like magic, if we get dpark speculative decoding versions of these models it will further improve the throughput
- schmuhblaster 3mo agoI've worked extensively with the slightly less able cousin, the 35B A3B model and tuned my own harness around making it work well with local or non-sota models. The results are quite promising [0], if one sticks to a plan-execute approach. After a bit of fiddling with llama.cpp I was able to get it to work through a small change on a real codebase from work on a 32GB M5 (typical python FastAPI backend, so nothing out of the ordinary). While that's somewhat encouraging, the whole local experience was still far from pleasant with all the noise and heat. [0] https://deepclause.substack.com/p/how-to-make-small-models-punch-way https://deepclause.substack.com/p/how-to-make-small-models-p...
- zbendefy 3mo agoWhat harness are you using?
- schmuhblaster 3mo agoIt's my own (slightly idiosyncractic ;-) harness: https://github.com/deepclause/deepclause-sdk https://github.com/deepclause/deepclause-sdk
- modgate 3mo ago[flagged]
- imrehg 3mo agoI'm having a decently good time time with `qwen3.6-35b-a3b-mtp` (unsloth's multi-token prediction version) and and `qwen-agentworld-35b-a3b`. On a 2021 M1 Pro (32GB RAM) I can get either of them as `IQ4_NL` quantized models (the first with reduced context, around 160k; the second can do the whole 264k with RAM left over), running something like 30tokens/s. On a Framework 13 AMD AI HX370 it can use the same, but both on Q8_0 quantization, full context window, parallelism. Speed is just ~15tokens/s so slower, but definitely smarter than the lower quantized siblings. Both of them are good developer partners for an engineer who wants more of a second pair of eyes and a rubber duck, rather than a model to just do everything for them. Pretty good for my brain dumping, some commit reviews, sanity checks, just always assume that every claim has to be checked and re-checked. The only problem is really the context loading, that's pretty slow (starts off around 300token/s on empty context, by the time we get to something like 70-80k which is just a bit of repo discovery, it can run around 80 prompt token/s or less, so there's always a lot more waiting around. Local tools need to bump all of their timeouts, and have to be mindful that there's unlikely to be really meaningful parallelism on these machines with local models. I'm still figuring out how to approach these things, though. Definitely better than glorified autocomplete or search tool (and too slow for the former, pretty decent for the latter). Their limited skill and performance make it more in line with other tools like my IDE or editors, that they are still in the "tools" compartment of my thinking, rather than "independent, cognitively active entities". Which feels like a good thing.
- nunodonato 3mo agowhat are you using agentworld for?
- imrehg 3mo agoFrom the Huggingface page, it is a fine-tuned version of the Qwen3.6-35B-A3B that is my alternative, and the benchmarks seems to be better. So I'm using as a "likely some quality gain over the other model, while the performance seems pretty much the same". I've done a couple of checks, and it seems very marginally better on some local benchmarks I'm running, but it's not super scientific evaluation.
- PeterStuer 3mo agoBeen running it on a 9950x3D with 96GB and a 4090. Speedwise it is fine. Bit while not completely useless, for software development it is unsurprisingly a dramatic downgrade from the Opus I use as my daily driver.
- grokkedit 3mo agoI've been using it with a couple of tools (like context7) as a documentation/helper, without giving it direct access to writing code, in marimo. it works great, albeit a little slow on my server (m1 max 64gb ram), at 8bit with omlx
- Go7hic 3mo agogoat
- meta-level 3mo agowhy does everyone imply you need a $10k laptop which then starts burning when you run Qwen 3.6? Get any other system with enough VRAM for a third of the price. Framework Desktop (Strix Halo 128GB) still costs under 4k nowadays, is nearly silent even on 100% GPU + CPU. (also it gets only slightly 'warm', but with a desktop you don't care anyway, I guess).
- paintbox 3mo agoBut how will I signal my status to other people then? On a serious note, I run my models on desktop pc, simple api and i can use them wherever whenever.
- hendry 3mo ago[dead]
- LoganDark 3mo agoI see OpenCode mentioned in the article, and I would strongly warn against using it for local development because it disrespects caching (the content of the first turn / system prompt is NOT stable). I use Pi which works much better.
- Frankybeatz 3mo ago[flagged]
- letmetweakit 3mo agoAny chance to run this on a RTX 3090 and 64GB of regular RAM with decent context size?
- amlord 3mo agoTried looking at it, but needs a much beefier machine than I have RN. Hopefully we're looking at a future where local models become more & more realistic to use for reducing remote AOI spend.
- ermantrout 3mo ago[flagged]
- _tyiueojdfg4 3mo agoMy personal experience below: I ran into some small problems with codex during setup and, for a few reasons, did not want to set up a cli shell with them at the time. Since I was not doing anything really serious, but just exploring a half-baked idea for an android app, I ran qwen in lms and connected it to android studio. None of the mini projects that I have attempted ( more granular call control, silly html scrolling game, music play app ) were one shots despite carefully preparing the prompt ahead of time. Admittedly, some of it may have something to do with android studio, but I did not try it with google account yet. All took between an hour to four to generate ( prep, initial run, testing, iteration and so on ). If it helps, miniforum AI MAX 395. I am not saying it is bad. Quite the opposite, but you want to be aware of the limitations though and plan around those.
- alper 3mo agoI have a fairly beefy M4/48G but I haven't been able to get any local model to behave anywhere near satisfactorily.
- Roark66 3mo agoOn dual rtx3090 it runs at 140tok/s with a short prompt... Not bad. Qwen 3.6 dense runs at 40tok/s
- yashthakker 3mo ago[dead]
- aichi 3mo agoWhat model fits on 36GB RAM mac?
- john-frandsen 3mo ago[flagged]
- love0972 3mo agoWhich one is actually better between Qwen and DeepSeek, and which one costs less?
- shekharupadhaya 3mo agoCan it work on a base mac mini 4 as a server to serve api for ai client mobile app?