17 ms·
Ask HN: Has anyone replaced Claude/GPT with a local model for daily coding?
Has anyone here fully swapped Claude/GPT for a local model as their main coding tool, not just for side experiments? If so, please share your setup and performance (e.g tok/s)
- v3ss0n 4mo agoYes qwen 3.5 122b+ dgx is working wonders and I ko longer subscribed to any cloud api now. I will post a project which I accomplished in 9 days of long horizons running.
- iluvcommunism 4mo ago[dead]
- tumetab1 4mo agoNot yet, tried Gemma 4 on an Apple M4 but the tok/s is significant lower than the cloud offering. Also,the lack of enterprise tooling to help selected an appropriate model and tooling to run a local LLM does not help.
- arjie 4mo agoNot “local” and not interactive coding but sharing since it might be helpful. I have 2x RTX Pro 6000 Blackwell running DeepSeek V4 Flash. I get 160 tok/s raw but it’s a reasoning model. For my use case, I have it auto-write code and another system auto-review the code. I occasionally use it with pi to write some code and it’s blazing fast but it’s mostly habit that keeps me with CC and Codex.
- leptons 4mo agoHave you measured your electricity consumption for this rig? I have to wonder how much it would cost you per month.
- ux266478 4mo agoNot nearly as much as you might think. 1.2kw where I live translates to about $0.12/hr, and that's when running full clip. If you have a decent solar hookup, it's small fraction on a sunny day. The expensive part is the upfront hardware cost and the electrical system upgrade you'll need to give your house.
- leptons 4mo agoI'm paying about $0.19/hr and using half that power just for a large spinning RAID, running some VMs and security cameras. And I'm reconsidering my digital extravagance because of the electric bill. You probably make way more money than I do.
- mtone 4mo agoHere's a DeepSeek-V4-Flash benchmark on 2X RTX Pro 6000: - Prefill: ~10K tok/s - Decode: 190 | 375 | 980 tok/s (for 1 | 4 | 16 concurrent requests) - GPU power draw during benchmark: Average: 585W | Max: 849W | Limit: 1200W with undervolt. Idle PC is 125W. I've asked it to calculate the following considering a realistic blend of cached prompts and decode for agentic dev scenario. Electricity-only (@ USD $0.08/kWh) Usage | IN price | OUT price | Monthly cost Concurrency=1 | $0.040/M | $0.080/M | $8.65 to $38.88 (5% to 100% active) Concurrency=4 | $0.024/M | $0.044/M | up to $48.67 (cheaper per token but higher power draw) Total cost of ownership over 3 years is electricity + USD $20K (pre-hike pricing). In a production scenario, how much would I have to charge my users to break even, aiming for 4 concurrent requests 24/7? A) Breakeven API pricing (est. 2B IN + 1B OUT throughput/month): IN price OUT price Self-hosted $0.121/M $0.363/M OpenRouter (budget) $0.098/M $0.196/M OpenRouter (DeepSeek) $0.140/M $0.280/M B) Breakeven subscription (users active ~1.5h/day): 1 user: $563/mo (oh, hai) 25 users: $23/mo 100 users: $6/mo
- arjie 4mo agoVouched your comment. Very cool. What are you running on to get 190 tok/s? I get 400 tok/s at c=4 but c=1 is slower than you.
- CamperBob2 4mo agoNot OP, but I am seeing up to 260 tokens/second output at c=1 with the recipe at https://github.com/local-inference-lab/rtx6kpro/blob/master/models/ds4-flash-v4.md https://github.com/local-inference-lab/rtx6kpro/blob/master/... using 4x 6k cards. Average is more like 200. There may be a way to get the 2-bit quantized version running even faster on a pair of them.
- arjie 4mo agoThank you. Useful to know. Clipped on top by reduce, I assume.
- akersten 4mo ago> I have 2x RTX Pro 6000 Blackwell Where did you find/order these? All the sites I can find are either out of stock, only sell to businesses, or are otherwise sketchy...
- arjie 4mo agoI run a small business (https://technologybrother.com https://technologybrother.com) that runs a few small SaaS so I ordered the GPUs through corporate sales. If the barrier is getting an LLC, those are relatively cheap. The nice thing is that if you've got a legitimate business with use for GPUs you can get into the Nvidia Inception Program which has a pretty solid discount.
- zackify 4mo agoMicrocenter is the easiest place but almost any vendor will sell to you after you email them and if you have an LLC
- CamperBob2 4mo agoCentral Computer is a good source in my experience: https://www.centralcomputer.com/all-products/ai-components/ai-gpus.html?p=1&product_list_limit=60&product_list_order=high_to_low https://www.centralcomputer.com/all-products/ai-components/a... No affiliation, I've just ordered from them a few times.
- alexellisuk 4mo agoWhat quant?
- arjie 4mo agofp8
- HappySweeney 4mo agoI have an optane and lots of ram, so I tried full-fat models for writing some function overnight, as I get about 0.7 t/s. My current go-to test is to update a scalar function to transpose a bit-matrix to one using avx512. the cloud models all play with that like its nothing. Kimi 2.6 and GLM 5.1 both failed miserably.
- kertoip_1 4mo agoJust attach OpenRouter to your coding agent tool and try yourself. All relevant open weight models are there. Every person have different needs and expectations
- christkv 4mo agoWaiting for this https://github.com/antirez/ds4 https://github.com/antirez/ds4 to stabilize for strix halo.
- acc_297 4mo agoI've been wondering lately if it would help to take a medium sized model and either in cloud or some local setup actually do Reinforcement Learning from Human Feedback (RLHF) on every prompt as a chore - I don't know if trying to manually finetune a model to your use habits would ruin it or help - ideally if you were diligent you could get rid of some of the ticks that make models for the general public difficult to work with e.g. overly sycophantic, overly verbose, annoying tendency to explain via analogies but perhaps one individuals prompt feedback just isn't going to ever be enough I'm not sure how much you need (I know people working at big companies that have purchased in-house agents fine-tuned on internal documents etc.. and apparently these end up with bizarre behaviours not necessarily more helpful than the standard models) I'd like to be able to essentially edit every response given by an agent and then finetune on the difference between what it produced and how I edited the text. Personally I would just remove a lot of the adjectives and try to distill the responses to core responses but I worry based on some of the work done by Owain Evans and other alignment researchers that this can sometimes push agents into tricky-to-predict tendancies.
- rolisz 4mo agoI'm interested in trying something similar. I was thinking to do this for my OpenClaw agent. About Owain Evans work: I think he did SFT. On Twitter someone was saying that RL is not as susceptible to what he showed. I'd like to try that
- htrp 4mo agoCursor is doing that (i think with Fireworks as their provider) https://cursor.com/blog/real-time-rl-for-composer https://cursor.com/blog/real-time-rl-for-composer
- dude250711 4mo agoYes, running a local model on a natural wetware substrate here. Recommended setup: plenty of nutrients, some caffeine and a quiet environment. Performance - not currently measured in tokens: roughly average.
- HPsquared 4mo agoI personally get about 50 tokens per hour.
- jasongill 4mo agoI have been running this stack since well before Claude Code became popular. It works OK but I've found it to be very slow; and despite having a big context window, it seems to lose track of what it's working on and goes down a rabbit hole (or just wastes tokens trying to use the web browser) for hours and is hard to get back on track. I even tried spinning up two sub-agents but even after years of trying to prompt them, they are almost useless in terms of coding ability, so that is looking to be a waste of spending at least so far but maybe the model will improve as time goes on.
- bananadonkey 4mo agoMy sub agent has been looping for almost 10 years at this point and has so far written 0 lines of code. Definitely won't be investing in another...
- phlhar 4mo ago[dead]
- K0balt 4mo agoPretty good results with qwen 3.6 27b dense. I’d say it’s about equal to (Claude) haiku 4.5 maybe sonnet depending on the task.
- kandros 4mo agoI’d rather ask my butcher than Haiku for coding tasks
- papichulo4 4mo agoAgreed on this. Anthropic has now changed the verbiage on the definitions of the models under `/model` to say that Opus is for everyday usage, and Sonnet is for routine tasks. There's apparently a reason Sonnet and Haiku have been left in previous version #s. Still encouraging, though, that things are catching up. We can't expect $20k local setups to match $20bn compute clusters.
- deleted 4mo ago[deleted]
- K0balt 4mo agoI’d say when qwen works it works like sonnet, when it fails it fails like haiku. So it’s less consistent but works pretty well, I guess? It’s still overall pretty useful for a lot of stuff, and I can run it directly on my MacBook. Once you get an idea of what it can and can’t bite off, it’s pretty easy to break things into chunks it will handle reliably with grace. But I still like to have access to SOTA models for review. Also you can have a SOTA model write a development plan that is basically a bunch of prompts to generate each part, then have the local model follow the plan. I should mention not to run it at less than q6, I prefer q8.
- kadoban 4mo agoWhat tool do you use to drive things for you, out of curiosity?
- 4mo ago
- Razengan 4mo agoRelated: Are there any viable distributed AI models? Like how we've had SETI at Home, Folding at Home, BitTorrent etc. People are clearly willing to donate their computer resources to distributed projects. Maybe in a dAI network anyone could submit content for training on, and each user running a "node" could have their own custom private conditions on which type of content to accept for training or inference. Like someone who dislikes anime could say "never accept anime related content or queries" so their node would basically opt-out from any data or questions about anime.
- joshuamoyers 4mo agoI think it'd be very hard to achieve viable tokens/s or get arithmetic intensity to be high enough in general, since many cases in existing training and inference are memory bandwidth limited. Definitely seems possible to conceptually have a slow pipeline that is distributed though.
- deleted 4mo ago[deleted]
- SimianSci 4mo agoThis is unlikely to happen in any meaningful fashion for quite some time. (TLDR; Distributed compute for models will require hardware at a level only really possible with data-centers at the moment.) Token generation operates at such a scale to demand enough from a single GPU as it will often saturate the bandwidth capabilities of consumer grade interconnects like PCIe. Which fundamentally implies that distributing a model's compute across vast distances is too much of a challenge without significant infrastructure. To give an example, When we split a model's compute between two seperate cards on a single workstation, this doesnt mean we end up with 2x the compute bandwidth for a model. Instead the increase becomes something small like 20% depending on model, because the inconnects (PCIe on consumer hardware) will quickly become so saturated with data being copied between the two GPUs so as to become a bottleneck. And remember that this is something that happens locally with PCIe, which (depending on generation) will cap out at around 20-35 GB/s depending on the generation of motherboard. Model performance is very much tied to having the fastest and highest bandwidth single card available so as to keep data transfer operations to a minimum as the sheer volume of data necessary for the model to run is immense. I simply cant imagine how slow and unusable a model would be if the copy operations necessary for its compute needed to be performed over unreliable network speeds where there will be significant performance loss as network speeds are not reliably distributed across the globe, and their unreliable nature would demand increased overhead due to data verification. The dream of distributed AI is a ways off.
- _davide_ 4mo agoi used to mix remote and local minimax 2.7(q3) on my strix halo, it run at 30 tg and 220 tokens pp... it was a bit painful slow, but it was a good feeling i could stay offline. unfortunately m3 which is at opus .8 levels is 460b parameters and doesn't even fit in 128gb of memory, let alone a big context. strix halo feels like a toy for ai purposes. https://kyuz0.github.io/amd-strix-halo-toolboxes/ https://kyuz0.github.io/amd-strix-halo-toolboxes/
- sosodev 4mo agoMy strix halo board is feeling more useful and less toylike with the recent performance gains combined from MTP, better quantization, and generalized performance improvements across the stack. For example, I can run Unsloth's Gemma4-31B 4-bit QAT model with around 30tg and 200pp. I don't find that to be too slow at all. Particularly because it's nearly full accuracy and good enough for a lot of different stuff I throw at it. I think it also helps that I'm using my machine to do home server stuff. It excels at all of the traditional workloads. Then I can lean on the AI to help with automation here and there. I find it deeply satisfying.
- _davide_ 4mo agoyou can absolutely use it for some workloads, but as soon as you have some extra complexity for a big repo it'll take forever and the economics are so silly to the point that the electricity bill would be comparable to a subscription. I love having the possibility of running things locally if some random dude decide to pull them plug, and give me solice the fact that i can have 100% private inference, but as the main driver during the day? shoot me
- sosodev 4mo agoMeh. My server can run these models for neglible power draw (like ~130W fully maxed out). That's with ~30 tok/s which isn't that bad. I do agree that they're still nowhere near as good as the frontier models though. I do lean on those when I need to get something done with better quality or at a faster speed. I've also been using Deepseek V4 pro/flash for some work stuff and I do find them to be much closer to frontier capability. I may try running flash at home soon for very patient edits. :)
- ryandrake 4mo agoAlways a bit disappointed in the details in these kinds of threads. When you do get answers, they're never specific enough to try out on your own. It'll be something like "I use Qwen 3.5 and get great results!" OK but what quantization are you using? What llama parameters? What context size? What GPU are you running it on, and how much VRAM does it have? Are you hosting it on a separate box, or running it locally on your dev machine? What coding agent tool are you using, and how is it configured / hooked up to the model?
- riazrizvi 4mo agoAll you get here is some market signal from 1 or 2 posts if you already know how to do it. Most of these responses are garbage.
- porkloin 4mo agoI have good results with this setup: Hardware: - GPU: AMD 7900xtx, 24gb vram - CPU: AMD 5950x, AM4 - RAM: 64gb DDR4 3600 Software: - OS: Bazzite (atomic fedora - this machine is running Steam "big picture" mode on my TV when not in use for LLM tasks) - Virtualization: Podman Quadlets, which allows me to run container images as managed systemd units - Network: tailscale - Inference: llama.cpp vulkan (better performance than ROCM, though I'm keeping an eye on it in the future) - LLM API surface: llama-swap (running as a podman quadlet exposed via tailscale svc) allows running multiple models on a single endpoint. - Web/Chat Access: open-webui (running as podman quadlet exposed via tailscale svc) allows me to access any of the models I'm using for coding harness access for chat/general purpose queries via web browser. I also have the "conduit" app for my iPhone that allows me to hit the same models from my phone. Models: - Qwen3.6-27B-MTP-UD-Q4_K_XL.gguf - Unsloth Q4 quant of the qwen 3.6 27B model weights, with MTP enabled. MTP is important as it improves the speed the model can run at. - Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf - Unsloth Q4 quant of 35B-A3B. Not MTP right now because I was having some issues with it? - gemma-4-26B-A4B-it-UD-Q4_K_XL.gguf - Gemma 4, which I use sometimes via open-webui instead of Qwen, but I generally think Qwen does a better job Flags (specific for Qwen 27b, since that's primary model): - `-ngl 99` offload all layers to GPU - `-c 80000` 80K context window. I'd like this to be higher, but since my GPU also has to run the desktop session for the machine, I need to leave some VRAM overhead to keep the desktop from OOM-ing - `-np 1` single slot (no parallel request handling) - `--no-context-shift` error instead of silently sliding the context window when full - `--cache-reuse 256` reuse cached prefix in chunks of 256 tokens (prompt cache) - `-b 2048` logical batch size (tokens per submission) - `-ub 1024` physical micro-batch (per GPU pass) - `--cache-type-k q8_0 --cache-type-v q8_0` symmetric 8-bit K/V cache. Q8 is as low as I've been able to go without getting some issues with tool calling - `-fa on` flash attention - `--spec-type draft-mtp` use the model's built-in MTP as the draft model - `--spec-draft-n-max 3` propose up to 3 draft tokens per step - `--spec-draft-n-min 0` allow zero drafts if confidence is low - `--spec-draft-type-k q8_0 --spec-draft-type-v q8_0` KV quant for the draft path - `--reasoning-format deepseek` parse <think> blocks in proper format - `--chat-template-kwargs '{"enable_thinking": true}'` turns on Qwen's thinking mode on by default (clients can override) - `--jinja` use the GGUF's Jinja chat template - `--temp 0.6` moderate randomness (Qwen recommended value for coding) - `--top-p 0.95` nucleus sampling (Qwen recommended value for coding) - `--top-k 20` top-20 candidates (Qwen recommended value for coding) - `--min-p 0.0 disabled (Qwen recommended value for coding) Performance (27b, primary model): - ~65t/s for token generation - ~600 t/s for prompt processing. - If these numbers don't mean much to you, perceptually this feels about on-par with cloud model speed, maybe slightly faster. - ~30s cold start when swapping from a different model or starting up session from idle via llama-swap. I have llama-swap set up to unload the model after 10 min of idle, because I sometimes use this machine for gaming as well. A little annoying, but a small price to pay to be able to use the machine for other stuff (gaming) when I'm not using it with coding tasks. CLI/Harness: - Crush harness (https://github.com/charmbracelet/crush https://github.com/charmbracelet/crush) less feature rich than Claude Code, but with a smaller system prompt and better built-in LSP support. I point it at the tailnet DNS (https://llama.<tailnet>:<port>) - Headroom (https://github.com/chopratejas/headroom https://github.com/chopratejas/headroom) to maximize the 80k context window - Exa MCP for web search (https://exa.ai/ https://exa.ai/) this alone makes the model far more useable. It's shocking how often the official claude code or codex harness get botblocked on web fetches, and the results of a good web fetch can be the difference between a good turn and a bad turn. A lot of people get hung up on whether Qwen 3.x models are "as smart as" some parallel Anthropic model. Most people seem to agree it's somewhere between Haiku 4.5 and Sonnet 4.5. Personally, I think the biggest thing that makes the Qwen 3.x series of models _feel_ good to use for coding workflows is that its the first time that tool calling actually works consistently on local models. If tool calling is busted even 5% of the time, it can totally ruin the flow. I think that's also why people tend to say the "harness is more important than the model" or whatever. I have a few other models set up but 27B with MTP is the best compromise of speed and quality that I've found. This setup works well enough for me that I dropped my personal Claude Code subscription. At work I'm still using frontier models, but personally I don't feel like I need that much power for anything I work on in my personal life. I'm "lucky" that I made the random financially unwise choice to buy a 7900XTX in late 2022 for $1k as a gaming card. I had no clue it would actually be a pretty decent LLM card 3-4 years later. Edit: sorry for the horrible formatting, I always forget that HN doesn't actually do markdown :(
- anonymousiam 4mo agoThis was posted shortly after your Ask HN post: My Homelab AI Dev Platform https://news.ycombinator.com/item?id=48542433 https://news.ycombinator.com/item?id=48542433
- ecshafer 4mo agoI work with a few models on servers, so not local, but self hosted with ollama. gemma-4, glm 4.7 flash, and qwen 3.6. glm is the best at coding agentically. But I still don't think any of them reach the levels of gpt 5.5 or opus 4.8.
- system2 4mo agoUntil I can buy an 80GB VRAM GPU, I won't attempt to do it. A local LLM is always missing something that needs a bigger model.
- ColonelPhantom 4mo agoWhich model class requires an 80 GB VRAM GPU? From my perspective, popular models seem to be either in the ~30B range (Qwen3.6, Gemma 4), while the larger models (MiniMax, MiMo, StepFun, Deepseek) are in the multiple hundreds of billions parameters, for which 80 GB is simply too small. You can just about reach the lower end of the latter category with a 128GB machine like a DGX Spark, Framework Desktop, or M5 Max, though those are usually not super fast. For the former category, you can easily run them fast with something like a 3090 or 5090, hell, probably even a 5060 Ti.
- CamperBob2 4mo agoThis is true. There's not much point in buying only one RTX 6000. You need at least two to run anything interesting that you couldn't run on a 5090. And you can imagine where it goes from there.
- system2 4mo agoVideo models.
- mitchell_h 4mo agoTried. The context windows just weren't big enough.
- deadbabe 4mo agoPrompt more directly instead of open ended.
- lysace 4mo agoGot a similar result (my RTX 4070 only has 12 GB). I'm curious about whether 24/32 GB meaningfully improves this enough to make it useful.
- tobyhinloopen 4mo agoTry it on RAM and CPU. It’s slower but you can run them.
- lysace 4mo agoGood idea for evaluating the models, thanks.
- coder543 4mo agoQwen3.6-27B supports a 1 million token context window. Of course, you have to have the right hardware to be able to run with a context window like that, as it takes about 100GB of memory on my DGX Spark to do that with full f16 KV cache on the q4_k_xl model.
- nfrankel 4mo agoI tried. It works in theory: https://blog.frankel.ch/tokensparsamkeit-coding-assistants/#local-models https://blog.frankel.ch/tokensparsamkeit-coding-assistants/#... Results depend on the model, of course, and your computer is the limit. Mine wasn't up to the task, unfortunately.
- codinhood 4mo agoI don't think you're going to get many "true" answers to this. The opportunity cost of not using the latest and best models is just too much right now. Every month I research this and come to the same conclusion: the time, effort, and cost required to get local models (and the coding tools around them) to perform even close to Claude Code with sonnet/opus just not worth it right now. If it was, it would be distributive enough to be in the news. Not that I'm discounting someone hasn't already solved this, just trying to Occam razor my way out of diving too deep down rabbit holes.
- jrm4 4mo agoBut you're pretty much measuring opportunity cost in tokens per second, no? I think it strongly remains to be seen whether e.g. tokens per second (multiplied or whatever by percieved quality of private model) actually means "better or more useful output." I strongly suspect it does not. (though I also strongly suspect this will be very difficult to measure because the incentive to lie about metrics here will be so strong.)
- codinhood 4mo agoIf you’re arguing that model metrics don’t necessarily translate into useful output, I agree. That’s not how I measure the success of a mode and not really the point I'm trying to make. I try to set things up and test it on my actual projects. What I’m saying is that if local models were actually comparable to Claude Code in practice, we wouldn’t be having threads like this. It would be obvious to the people using them, and it would be massively disruptive. Why would individuals and companies pay hundreds or thousands for Claude Code if they could run something locally and consistently get similar results? Every month I revisit the local ecosystem hoping the answer has changed. So far, my experience has been that it hasn’t.
- jrm4 4mo agoHaving, e.g. seen Microsoft maintain a monopoly for well over a decade, there's nothing in my experience that suggests that "quality always beats hype" is remotely true. It's entirely possible Claude is just winning the hype game.
- SkitterKherpi 4mo agoIt has so far been the kind of thing that always feels like the next version of the local models would be the one that is just good enough.
- Lwerewolf 4mo agombp16 m5 max 128gb, antirez/ds4, deepseekv4-flash. Works well for relatively dense (let's say <20k LoC per project) C codebases that are essentially a bunch of custom specialized stores, http servers, network infra, media transformers, etc. Runs through Pi with a custom prompt (basically "don't speculate blindly, isolate things, make them traceable and measurable, then verify") and behind a pretty restrictive bwrap setup - RO bind everything other than ~/.pi, cdw and a separate tmpfs, unshare almost everything other than the network - for which I use a network namespace that only allows tcp connections to a specific ip and port (i.e the inference mac) - i.e. netns exec into bwrap. Can't compare it to SOTA or higher-requirements models on what I work on - policy. That said, on a bunch of test pieces - it obviously isn't gpt-5.5, it definitely lags behind k2.6/glm/ds4-pro, but it absolutely is usable. Of course, on such codebases, forget about one-shotting or trusting it blindly or anything of the sort - you ask it, guide it, restart the context from time to time to have a "fresh dice roll" and to keep the context small and clean, etc. Compared to anything smaller (incl. all the usual local qwen models) - on a test piece, it figured out that memfd and mmap were used for setting up a ring buffer with natural wraparound handling (double mapping the first page at the end) and didn't tell me "this is for sharing memory between processes" or some other BS. Performance as described in the tables in the readme here: https://github.com/antirez/ds4 https://github.com/antirez/ds4 ...with a bit less than half that at "low power" (30w). Both are usable.
- dada216 4mo agoLocal? No. Via opencode Go subscription using GLM mainly? Yes, I still use Gemini/Claude/GPT via api from openrouter for adjacent tasks, I would say 20$ per month max in api token costs. Disclaimer: I am a Linux infra/k8s guy, I write production code but it's mainly glue code and mainly in golang. Addendum: most value we get is from "document intelligence" and that's all Gemma and Qwen on H100/H200
- sosodev 4mo agoThe problem with this question is that it encompasses a huge spectrum of capabilities and expectations. If you can only run an 8B model and expect it to be good at vibe coding / one shotting things you're going to have a bad time. If you're able to run a model on the scale of ~30B, you can find that with a reasonably scoped and well defined task they do very well. I've found both Gemma4-31B and Qwen3.6-27B to be the best in this range at the moment. You can swap in the MoE models for faster inference, but they are noticeably worse at most tasks. They can one-shot / vibe code tasks with small scope, but still do much better with guidance. If you really want frontier-like capabilities, you'll probably need at least 128GB of memory and either huge compute or a lot of patience. Most people just don't have either the money or the patience to make these local models work. The patience required for local model usage goes far beyond just waiting for tokens though. It takes a lot of effort to get things configured and working properly for your workflow and hardware.
- argee 4mo agoI use Gemma 4 26B A4B on my Macbook (M4 Pro, 48 GB RAM) to study Rust (and ask other myriad questions). I don't trust it to do a good job in an IDE/harness to one-shot anything but the most trivial of changes. Still, it's fast and good enough that it could handle being a "co-pilot" on small to medium context tasks where you've got your hands on the wheel and your eyes on the road — and are driving under the speed limit. That's remarkable given where we were a couple of years ago. I don't think I'd be using AI to code at all if this weren't the case. (I don't want to feel stunted or stuck just from losing my internet connection.)
- user43928 4mo agoMy experience with smaller models, in this case specifically GPT 5.4 Mini, is that they cannot two-shot moving a 10-20 line code change to another file without modifying it and introducing bugs. I did not expect perfect reliability, but I thought they could at least get it right on the second attempt once you point out the difference. No such luck, it confidently tells you that now the code is the same, with yet another subtle bug added in the difference. I don't know what work one would need to do where these garbage-class models would be adequate. Maybe they can masquerade as competent for a few minutes, but in the end the results simply are not right. At best they are suitable for a smarter search or autocomplete, in my opinion.
- blurbleblurble 4mo agoMy experience is that it's not the models themselves that are limiting right now, it's the clunky alternative harnesses with weird missing features making for bad ergonomics around stuff like queue management, interruption, subagents, goals, etc.
- Insanity 4mo agoHeard good things about pi.dev but haven’t tried it. It might take care of some of those missing features you mentioned.
- bityard 4mo agopi.dev is more like an agent developer kit. It's basically a substrate upon which you spend hours/days/weeks building your own agents or coding framework. It's pretty much the neovim to claude's vscode.
- horsawlarway 4mo agoI mean - the base experience is just fine, with perfectly reasonable built in tools for file access and editing, plus bash. But yes - it expands a lot if you're willing to play with it. I'd actually say the vscode comparison is wrong, because vscode is very much "bring your own extension" in the same way that Pi is. While Claude is much more "visual studio" vibes. It's thick, it's opinionated, and it's absolutely not something you can really customize, but it can feel slick for supported workflows.
- horsawlarway 4mo agoPi is decent. I've used the cli agents for claude, cursor, and pi, plus several custom harnesses I've written myself from time to time as experiments (and I guess technically gastown, if we're calling that a harness). Pi is... just fine. It does what I need it to, has a decent selection of tooling out of the box, integrates nicely with other tools, and generally gets out of my way enough that I don't think about it much anymore. If you can run ~30b models at decent speeds, I think most folks would be pleasantly surprised at how capable they are with pi. Tack on some of the extensions (ex https://pi.dev/packages/pi-mcp-adapter?name=mcp https://pi.dev/packages/pi-mcp-adapter?name=mcp and https://pi.dev/packages/pi-web-access?name=search https://pi.dev/packages/pi-web-access?name=search) and I get web tooling (ex - perplexity search), access to mcps to do things like drive chrome (https://browsermcp.io/ https://browsermcp.io/) or firefox (https://github.com/mozilla/firefox-devtools-mcp https://github.com/mozilla/firefox-devtools-mcp) It's fine. Is it as good as a subsidized top tier model? Nope. Is it free and still very capable? Yup. And personally, I've been having a LOT of fun with the pi sdk (https://pi.dev/docs/latest/sdk https://pi.dev/docs/latest/sdk) Which is something that all the other providers charge you api access rates for (ex - thousands a month).
- temilson 4mo ago[flagged]
- dabinat 4mo agoThere’s evidence that combining models can achieve frontier-level performance (e.g. OpenRouter Fusion). I’m wondering if that’s the more realistic option: combine Opus with a local model to save on token costs.
- rvnx 4mo agoI start to believe that adding more and more and more and more and more thinking tokens is the hack that works (this is what gave birth to Fable)
- utopiah 4mo agoWhy would you not think that? It seems pretty intuitive that pouring more resources into a problem (more GPU, bigger GPUs with more VRAM, bigger datasets, better curated datasets, more efficient ways to train, more efficient way to run inference, etc) then running the result for a longer time, with more layers of verification (running in VMs, model fusion comparing multiple models, having harnesses with testing) will at least lead to marginally better results. Is it worth it and at what pace will it keep on improving are different questions but I have little doubt that if the industry keep on pouring resources, sure more "works".
- pierotofy 4mo agoYes. Llama.cpp + Qwen3.6-35b (MTP) + OpenCode is quite capable and runs on a single RTX 3090 and is faster than most cloud models. Quality is like running edge models from 8-12 months ago. Setup details at https://github.com/pierotofy/LocalCodingLLM/ https://github.com/pierotofy/LocalCodingLLM/
- atomicnumber3 4mo agoSame. I have no desire to use Claude at all anymore.
- pierotofy 4mo agoYep. Screw Anthropic, CloseAI and all other rent seekers in this space.
- akulbe 4mo agoI have an M2 Max MBP with 96GB of RAM. What models and setup would you use for this kind of configuration?
- monirmamoun 4mo agodownload LM Studio to play with, and it will let you search for models... try Qwen3.6-35B-A3B at 4,5 or 6 bits (6 bit XL is near perfect) and use pi coder or another harness to access it... you can also try Unsloth studio and try same model to start. LM Studio slighter easier to use, Unsloth probably better quality. Neither one is super great quality by the way (meaning: they crash or act weirdly too often to be full production solutions, but can work for local coding). ONCE YOU DOWNLOAD EITHER APP... it will let you search huggingface for the models. Just type qwen to start looking and ... start messing around. And you connect the pi coder harness using the http interface that LM Studio and Unsloth offer to the engine API, so make sure you figure out that url and turn it on... something like 127.0.0.1:1234/api would be a typical IP (localhost) and port (1234 is used by LM Studio)
- jacobgold 4mo ago
- wuschel 4mo agoI would like to know whether someone was able to use lower tier models for activities other than coding e.g. a limited version of a personal note manager - and what the hardware requirements in RAM for these models were.
- fortyseven 4mo agoI use Pi and Qwen 3.6 27b locally on a 4090 for all my personal projects. I still use Claude for day job work since they pay for it, and my employer expects me to use it. I rarely touch it otherwise.
- cheekygeeky 4mo agoOur software dev (smartest guy I ever met) is using OpenCode and Tmux with Open Source models. He says the DeepSeek is his model of choice for coding (he call's it "pretty GOOD". He's running two 3090s on an i9 with 128GB RAM. https://www.msn.com/en-us/news/technology/china-s-open-deepseek-v4-now-scores-within-a-fraction-of-a-point-of-claude-on-a-key-coding-test-at-roughly-a-tenth-of-the-price/ar-AA252j22 https://www.msn.com/en-us/news/technology/china-s-open-deeps...
- cuttysnark 4mo agoI've had some success with local models by chaining "agents" together in a workflow. Each agent has a different prompt and uses a different ollama model based on what their role is. The project manager, schema agent(qwen3:14b), etc. doesn't use the same model as the coding agent (qwen2.5-coder:7b). Between each step is an orchestrator and with a Playwright task which attempts to surface errors to the agent who introduced the previous code block. Only error-free blocks are forwarded to the next workflow step. Probably the biggest improvement was including a backend-for-agents service definition which instructed the schema agent they were to only produce only a manifest based on the task, and to pass off that off to the next agent. In short, I split tasks up into many pieces by defining a workflow where agents are only allowed to do very specific things before their work is passed along. This keeps them grounded and capable while also creating places for me to intervene if a workflow was say 25% or 90% successful.
- pianopatrick 4mo agoI wish someone would do a benchmark and competition for this kind of work flow so we could figure out what works well. Like "Here's this consumer grade GPU. Using only this GPU but with whatever models and workflow you want, see how well you can do on xyz benchmark." Contestants would be given like 1 hour max and scored based on % of questions answered, % of questions correct and total time to finish. Like "The Local AI challenge"
- Curiositry 3mo agoTell me if you find this! I was thinking the exact same thing.
- sowbug 4mo agoHave you (or anyone else) tried letting agents compete? For example, give the same coding task to two models, or to the same model with a different seed, and have the reviewer choose the better result. Some think the human brain works similarly: thousands of mini-brain cortical columns, each with a slightly different take on the situation, voting in a majority-rules system.
- jmichaelson 4mo agoI am working on exactly this issue right now. My approach is that a highly optimized harness (pi.dev) with the right backing knowledgebase (a custom, self-updating wiki with lots of QC layers) can get close to most of my usage patterns for my Claude Max 20x subscription. I use Gemma 4 26B QAT served by a custom fork of llama.cpp, with 4-8 slots of 256k context at Q8. It's a very good model when the harness keeps it on rails. In an age of 1M context windows, 256k may seem small but it's been plenty for my work (scientific programming). A $20/month subscription to Ollama-cloud gets me good coverage of consults out to frontier models for difficult plans or debugging (again this is all woven into my highly customized pi install). I'm still optimizing it (with claude, to be clear), but my testing is very encouraging. I worry a lot about companies (and the government) controlling access to machine intelligence, so local is the way to go.
- horsawlarway 4mo agoFor personal use, yes. I replaced a $100/m subscription to claude in favor of running pi harness pointed at unsloth studio, using both qwen (unsloth/Qwen3.6-35B-A3B-MTP-GGUF) and gemma (unsloth/gemma-4-26B-A4B-it-GGUF) models, depending on my mood. I have a machine I built about 5 years ago with dual RTX3090s in it (I was going to build a new gaming machine anyways, and the llama release had just dropped so I tacked another used 3090 onto the build), and I get ~150tok/s on either of those models (at UD-Q4_K_XL quant) and can use the entire 300k context length without having to exit VRAM. To be very clear - it's not as good as claude. But it's free and not so much worse that it matters significantly. For my personal needs, free beats $100/m. I also have an openclaw instance pointed at the same inference server, and it's great for that (genuinely solid use-case for local models). Some example projects - Replacement launcher for android tvs (with usage monitoring and tracking for kids) - Custom admin portals for my k8s cluster services - Custom home assistant integrations/automations (recently some shelly devices for power monitoring and switching) - Grocery list management and meal planning (mostly via openclaw) - some custom workflows for 3d asset generation in comfyui. --- Long story short, if you're trying to make money via software... I'd probably still recommend using a paid provider. But the local models are very capable of cool stuff.
- gonzalohm 4mo agoDid you double the tokens per second by adding a second GPU or was the increase significantly less?
- mirekrusin 4mo agoYou’re adding extra gpu for more vram, not speed.
- horsawlarway 4mo agoNo real change in inference speed. It basically just allows me to slot in more context or a bigger model. A single RTX-3090 will do approximately the same tok/s, but it won't fit the entire 300k context in VRAM. Sometimes that matters, a lot of times it doesn't. On the speed front - MOE models are great. Biggest perf difference in modern models is the move to MOE architectures. I get very similar quality from the both the Gemma-4 31B dense model, and the Gemma-4 26B MOE model (both at Q4 quant) but the MOE version runs at ~3 times the speed (150tok/s vs 46tok/s).
- stymaar 4mo agoYes, Qwen3.6-35B-A3B on a Strix Halo 128GB (Bosgame M5). I have way too much VRAM forme such a model but Qwen never released the 122B version of Qwen3.6, which is the best class of model for my hardware. But at the same time my electricity bill is negligible, this is originally a laptop chip and it shows, it consumes almost nothing while idle and a little above 120W during prompt processing. And Qwen3.6 has been surprisingly effective for me, I still use Clause occasionally but only for like 10% of my needs which allows me to stay well under the quota even with the cheapest plan. Speed: ~800tps prompt processing and 50tps for token generation (with no speculative decoding).
- manmal 4mo agoHave you tried the 27B dense version? It’s way better for coding.
- hegdeezy 4mo agoI have tried locally but I find that the implicit breakeven is somewhere around 1 year of use given the high power costs where I live. Not really worth it but maybe if I move some day!
- BiraIgnacio 4mo agoI tried for a bit, with llama.cpp + Qwen + Mac Pro but the results were very poor (both quality and speed). I considered investing in better hardware but doing the math, it is cheaper for me to pay for DeepSeek (yeah, I know not everyone can do that).
- Kostic 4mo agoFor personal needs I connected VSCode with llama.cpp running Qwen 3.6 27B or Gemma 4 31B and it's good enough to cancel my cloud subscription. Qwen running on my 1st GPU at q4@176k context from 70 to 50 tok/s with MTP, pretty good for coding. Gemma on the other hand is using both GPUs, running q8@64k context, doing document sentiment analysis, summarization, proofreading and translating, at consistent 25 tok/s. Somewhat slow but usable for batched workflows. Might get some more once llama.cpp starts supporting MTP with tensor split mode. Still using frontier LLMs at dayjob since I'm not paying it and those are obviously better. Hopefully we'll have a Sonnet 4.6/Opus 4.5 level 30B model in a year or so. EDIT: Prompt processing starts from 800 t/s and drops to 400 t/s. In most cases my starting prompts are around 16k-24k of tokens and require from 60 to 90 seconds to be processed. Not great but acceptable.
- fitzn 4mo agoWhat extension do you use in vscode to connect it to local llama.cpp? Or do you auth with github copilot and then point to localhost? Or something else?
- khimaros 4mo agoi made this specifically for use with vscode/llama.cpp: https://github.com/khimaros/mortar https://github.com/khimaros/mortar
- Kostic 4mo agoAuth with Github Copilot and then point it to localhost[0]. Hopefully the auth to Copilot requirement will be dropped for local models at some point. Would love to use a fully open stack (VSCodium and everything) in the future. My config: ``` [ { "name": "http://127.0.0.1:8888/v1 http://127.0.0.1:8888/v1", "vendor": "customendpoint", "apiKey": "llama.cpp", "models": [ { "id": "gemma4-31b", "name": "Gemma 4 31B", "url": "http://127.0.0.1:8888/v1/chat/completions http://127.0.0.1:8888/v1/chat/completions", "toolCalling": true, "vision": true, "maxInputTokens": 65536, "maxOutputTokens": 8192 }, { "id": "qwen3.6-27b", "name": "Qwen 3.6 27B", "url": "http://127.0.0.1:8888/v1/chat/completions http://127.0.0.1:8888/v1/chat/completions", "toolCalling": true, "vision": true, "maxInputTokens": 180224, "maxOutputTokens": 8192 } ] } ] ``` [0] https://code.visualstudio.com/blogs/2025/10/22/bring-your-own-key https://code.visualstudio.com/blogs/2025/10/22/bring-your-ow...
- jwr 4mo agoI tried many, many times and I keep trying. But I just don't see this happening: those tiny models that we can run on our machines (I have an M4 Max Mac, so I can reasonably run qwen3.6-35b-a3b or gemma-4-26b-a4b-qat at this time) are NOWHERE near as smart as the huge monsters like Opus/Fable. Nowhere. I can see a lot of people deluding themselves. Sure, you can get the local models to generate plausibly-looking code for simple cases. But compared to how I solve complex design problems in a large codebase with Claude Code and Opus/Fable, this isn't worth my time.
- anubhav200 4mo agoYes, llama.cpp, qwen27b, 35b, claude code. Llama-cpp-manager for managing llama.cpp configs (https://github.com/anubhavgupta/llama-cpp-manager https://github.com/anubhavgupta/llama-cpp-manager)
- anubhavgupta 4mo agoMachine: CPU: intel 275hx GPU: Nvidia 5090 Mobile (24GB) RAM: 64GB
- anubhavgupta 4mo agoOne more thing, I also use it along with Whisper-NPU, a speech to text utility that runs on NPU of Intel 275hx and doesn't consumes any GPU resources.
- anubhavgupta 4mo agoWhisper-NPU (https://github.com/anubhavgupta/whisper-npu https://github.com/anubhavgupta/whisper-npu)
- boringg 4mo agoWill the AI labs always make sure there is at least a years worth of differential? I guess the underlying business premise is that each new release has a step function change that prevents this kind of behaviour..
- snoman 4mo agoIf the government is going to gate access to frontier models from here on out, even if new releases are a step function change… which they’re not… then it may be even more comparable to what’s available with a subscription.
- anubhav200 4mo agoYes, llama.cpp, qwen 27b and 35b, llama-cpp-manager for managing model configs.(https://github.com/anubhavgupta/llama-cpp-manager https://github.com/anubhavgupta/llama-cpp-manager)
- NetOpWibby 4mo agoI'm looking forward to having Claude Fable at home. THAT is when I'll THINK about replacing Claude (who knows what their next models will be capable of, Fable was damn good for the three days I had it).
- trueno 4mo agowe keep moving the goalposts on when we're gonna be happy with local. first it was sonnet at home as the good enough, then opus, now it's the mysterious leading model that runs on infrastructure we can't feasibly have at home
- NetOpWibby 4mo agoI don't know about "we" but for me, I've never been happy with any of the models to bother with learning how to run one at home. I've got a Turing Pi and a bunch of other gadgets just sitting in a box (well the former I've got running after owning for several years of non-use).
- bluejay2387 4mo agoAbout 90% of my coding is on Qwen 3.6 27b and Open Code with some custom skills and Semble. It is NOT as smart as CC or Codex but its enough to get most of my work done. I didn't set out to replace CC and Codex (I have an RTX 6000 so the TPS is faster than I care about, but the RTX 6000 was originally for other work). I only tried this just to see how close you could get to a frontier model for coding as an experiment, but it was good enough that I stuck with it. I still fall back to Codex for really complicated stuff and to polish UI's as that seems to be the weakest element to working in Qwen.This isn't a recommendation because I don't think most people have an RTX 6000 laying around and the cost would be many years of MAX CC or Codex subscriptions, but at least this seems possible. Maybe in a few more years it will even be practical. Other Notes: I have had to set the compact target to 75% on a 256k context window as once the conversation length goes about 100k I start seeing a drop in the quality and speed. This becomes very problematic after about 150k. I tried Qwen 3.5 122b too but it actually seems much worse at coding than 3.6 27b even though its much larger. Maybe because I am using a 4bit quant or maybe I just don't have it configured correctly? I know 3.6 is newer but I didn't expect it to out perform a model that is much larger from the prior generation. Gemma 4 31b is a good model for other tasks but at least my personal experience is that Qwen outperforms in coding. Nemotron Super 120b is great at a lot of stuff but it also seems to be not as good at coding as Qwen. This was very surprising to me.
- htrp 4mo agowhy 27b vs 35b? Is MoE that much worse for coding?
- electronsoup 4mo agoYeah MoE is a little worse for the same size, but you can often run bigger MoEs at respectable speeds even on cpu ram offload. The dense models really need to be 100% vram
- amarshall 4mo agoCan take the geometric mean of total and active parameters of MoE to get approximate equivalent quality to dense model params. So sqrt(35*10)≈18.7. The trade-off of MoE is that it is worse but faster for the same total size.
- AH4oFVbPT4f8 4mo agoOllama + Hermes on M5 Max 128GB using .NET using Qwen 3.6:35b-a3b as the primary model to do the work. I might use 27b to plan what to do.
- xeonax 4mo agoWhats .NET doing in between?
- AH4oFVbPT4f8 4mo agoSorry, I meant to say I was writing .NET C# with the setup
- mv4 4mo agoI've been using MiniMax M2.7 with vllm on my dual Nvidia Spark cluster. Slow (<20 tps) but functional for most of my use cases.
- cmrdporcupine 4mo agoI was just looking and it should be possible to run this one on 3bit quant on my single Spark? Maybe? Depending on context size? Assuming 3-bit doesn't totally lobotomize it.
- zaptheimpaler 4mo agoI tried gemma-4-26B-A4B just to see if it could help me read/sort my emails on a relatively under-powered setup (16GB VRAM + 32GB RAM) and it's not going well.. the model burns 24K tokens just on searching for the right tool and then dumps the email contents into context - i tried to get it to use code-mode to save context but the code-mode implementation can't save files so it was useless and im going to try to switch to "ssh-mode" into my devbox container. Still relatively new to this, so I'm probably doing something wrong
- anana_ 4mo agoPerhaps try a different model? Just from anecdotal experience, I find that the Gemma models smaller than 31B do not tool call as often as they should. Some of the benchmarks appear to back this up [0] Of course, a lot depends how you are using it (inference parameters, harness, prompting, etc.), but the model is quite important too. [0]: https://artificialanalysis.ai/models/open-source/small?models=qwen3-6-27b%2Cqwen3-6-35b-a3b%2Cgemma-4-31b%2Cqwen3-6-27b-non-reasoning%2Cqwen3-5-9b%2Cgemma-4-31b-non-reasoning%2Cqwen3-6-35b-a3b-non-reasoning%2Cgemma-4-26b-a4b%2Cqwen3-5-35b-a3b-non-reasoning%2Cgemma-4-12b#intelligence-evaluations https://artificialanalysis.ai/models/open-source/small?model...
- Rzor 4mo agoSo there was a problem with gemma 4 when it comes to tool calling that Google apparently fixed like 2 or 3 days ago. I remember reading something about this.
- gigatexal 4mo agoI tried to. I just couldn't get over how it made my otherwise whisper quiet M3 Max MacBook Pro 14 for the performance. The sweet spot has been adopting Claude Code to use the Chinese models. Deepseek V4 Pro is very, very good. But I am such a casual local user of AI that my 20/month Claude subscription is enough and I find myself using that more and more.
- devin 4mo agoAnyone here running a tinygrad?
- redox99 4mo agoModels that you can run at home (Like Qwen 35B) aren't remotely close to Opus or GPT 5.5. Not even close. The only open models that are in that neighbor are around 1T params, so forget about running at home. It's kind of like driving a shitbox. It can often drive you from A to B, and some people will try to convince you it's fine. It's not. There's no logical reason other than absolutely requiring the privacy, doing it for fun, or niche use cases like airplanes and so on. If you can't spend the insanely subsidized $20 for codex, you can use an API for chinese models which will run circles around these tiny models.
- pbasista 4mo ago> Models that you can run at home (Like Qwen 35B) aren't remotely close to Opus or GPT 5.5. Is that characterization based on some objective facts or benchmarks?
- redox99 4mo agoBased on private test prompts I've run through OpenRouter.
- kube-system 4mo agoYes, there aren't any 35B models that are beating frontier models at just about anything generalized
- xgulfie 4mo agoI don't need a Ferrari to get to work
- orangeisthe 4mo agoBut you need the best tools to do the job
- gr_norm 4mo agoYou need tools sufficient to do the job in an economical way, optimizing for both cost and quality. That is what 'best' means. We don't give every engineer all the resources under the sun, only what is appropriate. I suspect many will realize millions more dollars are being spent than needed to achieve the highest marginal productivity gains, and reallocate accordingly. Who wants more of their money going to developer tooling, rather than bonuses?
- Greenpants 4mo agoI have! I care about data privacy and LLMs being free. I'm using the Pi coding harness but containerized and sandboxed, to make sure it's running completely offline. On my Mac Studio with 128GB RAM (or MacBook with 36GB RAM) I'm using Qwen3.6 35b, with only 3b active parameters so that it runs really fast. I've done a complete redesign for my website's homepage and blog with Django + Wagtail. The latter is interesting, because Wagtail is a bit less well-known, so the agent, without giving it internet access, doesn't always know how to develop for Wagtail. I've used Qwen3.5 122b for when things get more complex. At 10b active parameters, it's significantly slower though. I've noticed a few things compared to large models like Claude. For starters, you really need to know what you're asking, and be precise; it doesn't do much thinking for you. Any assumptions left open, and it'll take the easiest route to reach the goal (e.g. CSS in HTML), often not the best in terms of architecture. It gets into loops quite often, and surprisingly often gets the edit tool call wrong, after which it will spend lots of thinking tokens and re-read files instead of retrying (despite the system prompt suggesting so). Comparing agentic Qwen3.6 35b to Claude Opus is like a junior with knowledge across the board, that you really need to guide, versus a senior that thinks with you on architecture. If Opus gives a 15x speedup, local and fully offline Qwen gives a 5x speedup. Which, given that it's completely free, is still mind-boggling to me :)
- GardenLetter27 4mo agoCould the harness not check for a failed tool call and pass it to a small model for correction without clogging up the main context?
- Greenpants 4mo agoI'm actually quite sure that directly retrying the tool call would often fix the edit-call already. But these models have been trained to "think" for a while for any problem solving, so they'll presume the problem of the edit is more fundamental and spend unnecessary tokens filling up the context. I'll experiment more with the effectiveness of AGENTS.md rules for local Pi agents. I feel like smaller (local) LLMs just lack in attentiveness to elements in the context window, like precise instructions, compared to e.g. Claude models.
- major505 4mo agoYes. I use Owen on my MacBook m1 (16gb) daily, running inside Ollama. Works well. Is not particularly fast, and I need to create a custom imagem that sets the temperature of the model to zero starting, so I don't get over creative with its bullshit, but it works reasonable week.
- Der_Einzige 4mo agoSecretly the problems many people have with agentic coding are related to poor choice of sampling settings, but the world will wait several more years before this is understood well. top_p and top_k are garbage but they are intentionally kept on purpose because subsequent methods enable coherent high temperature sampling, which is an absolute no go for alignment/safety reasons. The secret to actually good agentic outputs even with small models? Llamacpp has support for this little known sampler called "top-n sigma". You should use that, set it to 1 and set temperature to literally whatever you want (it could be infinity) and your model will just magically work to your maximum context window. That's because long context generation is a sampling problem.
- major505 4mo agointeresting. using Ollama I created a simple markdown file similar do a Dockerfile creating an agent called "boring as fuck". This is the content FROM qwen2.5-coder:7b PARAMETER temperature 0 SYSTEM "You are a senior software engineer focused in php and Node.js.Your responses should be strictily technicals, without poetry or prose, and focused in safety. IF you are working with legacy code, just apply the changes with the best syntax possible."
- cyanydeez 4mo agonever started. using wither qwne3-xoder-nezt or qwen3.6 35b if youre shoopping for a new pc, very easy to justify 128gb vram
- jmward01 4mo agoHas anyone been storing their cc sessions for future training data on their own models? I'd love to build a system that fine-tunes on cc sessions and a good first step is capturing my own sessions well.
- abidlabs 4mo agoYes! https://huggingface.co/changelog/agent-trace-viewer https://huggingface.co/changelog/agent-trace-viewer
- jmward01 4mo agoDidn't realize they did this. I have avoided pushing data to huggingface. This is all -deeply- private info and I haven't really reviewed their privacy policies and the like. I'll give them a look.
- xhinker2 4mo agoYes, I have. 1. Two RTX 3090s in Linux 22.04 2. Running Qwen3.6-27B Q6_K_XL GGUF 3. Using my own harness AZPal, I build myself, also wire it with Hermes Agent, works fine 4. Many times it solve problem that Codex can't solve https://medium.com/p/f237d575e861 https://medium.com/p/f237d575e861
- bravetraveler 4mo agoI'm largely 'all natural', any of my little LLM usage is local. 128G Strix system, a not-super-dense Qwen or Gemma variant will get 50-80 tok/s output. Not subscribing to Anthropic/OpenAI/etc even in the unlikely event these are the last local models released; simply not needed. Entirely fine without and in-model tool usage covers my currency concerns.
- jodoherty 4mo agoI use pi with an RTX Pro 6000 Blackwell to run Gemma 4 31b to do all my agentic coding. I find it useful. This side project highlights a similar approach to how I scope and tackle projects at work now: https://git.theodohertyfamily.com/wg-wrap.git/tree/README.md https://git.theodohertyfamily.com/wg-wrap.git/tree/README.md https://git.theodohertyfamily.com/wg-wrap.git/tree/CASE_STUDY.md https://git.theodohertyfamily.com/wg-wrap.git/tree/CASE_STUD... You have to apply a lot of careful architecture and TDD to your approach. Eliminate technical risk by tackling hard things early and wrapping them up in a simple, easy to use interface. I find I can get some projects done 2-3 times faster than if I wrote them by hand. It can also save about 5-10x time on mundane or broadly scoped projects by helping me consolidate and try out ideas very quickly. Setup-wise, I switch between vLLM using nvidia/Gemma-4-31B-IT-NVFP4 and llama.cpp using unsloth/gemma-4-31B-it-qat-GGUF with MTP. I throttle the GPU power usage to 400W. My current llama.cpp setup gets token generation rates between 60-150 t/s depending on MTP draft acceptance rates. Prefill is between 1500-4000 t/s depending on context length/depth.
- yesb 4mo agoTried out the wg-wrap tool, might come in handy from time to time. Neat that it was made with a local model. Some issues: 1. `wg-wrap healthcheck` was all green even though unprivileged user namespaces was restricted via AppArmor (Debian). That check doesn't seem to work 2. DNS doesn't work (no domains resolve) if the config lists multiple servers e.g. DNS = 1.1.1.1, 1.0.0.1 3. Peer endpoints don't support domain names, only IP addresses 4. Minor: the tool doesn't add an implied /32 cidr prefix for single ip configs (common from some VPN providers).
- jodoherty 4mo agoHey, thanks for the feedback! If I get some time to circle back, I'll be sure to incorporate these into some new tests and address them. I want to set up a qemu-system emulator based testing approach so I can incorporate things like AppArmor and SELinux into end to end tests that include different environment configurations. Part of that will be setting up software defined networking so I can have things like DNS and wireguard VPN servers in a box and then test and evaluate the wg-wrap behavior at the packet level.
- moezd 4mo agoNot yet. Without pure Apple game or decent GPUs, even with a lot of RAM and threads, all you get is about 30-50 tokens/second, and that's thinking turned off. Without these optimizations your model will have a field day with your MCPs, skills and agent descriptions and you will watch the paint dry before seeing the first output token. Local model serving means you have to fight for every token in your context window, which is quite opposite of what Claude/GPT/Copilot are pushing the industry towards.
- amarshall 4mo agoThinking doesn’t change output speed. Anthropic’s models are ~ 40–60 t/s median output speed.
- moezd 4mo agoDo you have access to Anthropic model weights to run them locally?
- amarshall 4mo agoNo, and having that is not required to know output speed nor the effect of thinking, so I don’t see the point in such a superfluous, indirect question. As for the question you’re likely asking: benchmarks that include speed across many models and providers available at various places e.g. https://artificialanalysis.ai/leaderboards/models https://artificialanalysis.ai/leaderboards/models
- pianopatrick 4mo agoI wish someone would do a benchmark and competition for this kind of work flow so we could figure out what works well. Like "Here's this consumer grade GPU. Using only this GPU but with whatever models and workflow you want, see how well you can do on xyz benchmark." Contestants would be given like 1 hour max and scored based on % of questions answered, % of questions correct and total time to finish. Like "The Local AI challenge"
- GodelNumbering 4mo agoAs someone that spends all day every day talking to LLMs, I'd say the OSS frontier models + a good harness is already a sufficient combo. For local deployments, we are missing one or two hardware generations (and may not get that soon since hardware companies are heavily favoring datacenter segment) to fully move to a local setup.
- bijowo1676 4mo agoOne of the interesting setups I saw is using expensive frontier models to write and update markdown for your app: specs, product requirements, architecture, etc but then use cheap/local model to implement the specs. Markdown is more effective at compressing information and fits the context window easier, than hundreds of source code files but this requires second and third passes, to smooth out the rough edges has anyone tried that?
- SupLockDef 4mo agoLocal isn't new for me. I am still coding my stuff, but Qwen3-coder:30b on my old rig with a gtx 1070 16gb RAM does wonders for me. I mostly use it as a google search if I forget a thing, or doing the boilerplates. I am using a mix of a non harness chat for the reply speed, and opencode / vim-ai for my boilerplates. $0.00 / month. That's the budget.
- jboss10 4mo agoHave you tried qwen3.6 or pi?
- SupLockDef 4mo ago3.6 is too slow on my old rig for some reasons, so I went back to qwen3-coder. I did try 3.6 on my main desktop. It was good, but I didn't see much differences than coder, so I am still using my old rig.
- grmnygrmny2 4mo agoJust sharing my $0.02 here - I have ethical objections to using OpenAI or Anthropic products so I was a reluctant adopter of LLMs at all. Local models address most, though not all, my moral objections so I’ve been using them for work and personal projects for about a month. The hardware I have (32gb Macs and a gaming PC with 10gb 3080) can only get me to Qwen3.6-35B-A3B at various quants but that’s enough (200-400 PP, 20-30 TG). It’s taken some time to learn how to best utilize it - some things take a bit of babysitting or direction - but it’s quite useful. Not having ever used CC I can’t compare but it’s been a great assistant or pair programmer for everything from embedded C++ to Vue. I wish I could run 27B as there have been moments when this model feels like it just can’t quite figure something out but those moments are quite rare. For a lot of tasks it’s a huge time saver and has proved super capable at digging into and fixing bugs given pretty vague instructions. I’m using Pi as my harness.
- eugmai86 4mo ago[flagged]
- KaiShips 4mo ago[flagged]
- whartung 4mo agoWill the inevitable M5 releases from Apple change this equation in any meaningful way? I'm waiting to swap out my last gen Intel iMac with a new M5 mini of some kind, with the eye to hopefully be able to run some models locally. I envision a mini (heh) arms race to simply swapping out an M(X-1) for an M(X) annually as this field shakes out.
- fransje26 4mo ago> Will the inevitable M5 releases from Apple change this equation in any meaningful way? No. Apple is also running out of RAM, so you will not have the RAM you need.
- bArray 4mo agoI'm in the middle of building my own based on LiquidAI/LFM2.5-1.2B-Instruct [1]. I run it on the CPU locally and get reasonable performance. I'm currently using it to solve small problems - but expanding it daily. [1] https://huggingface.co/LiquidAI/LFM2.5-1.2B-Instruct https://huggingface.co/LiquidAI/LFM2.5-1.2B-Instruct
- shironnnn_ 4mo agoI use SpecKit to create a very detailed plan with a high amount of specificity using paid Claude plan. Then I give it to local LLM (eg: Qwen / Gemma 4) via CLI. This is possible through usage of llm-mlx on Mac (or ollama on any machine given sufficient on hardware) which serve OpenAPI endpoints compatible for Aider (CLI) or Visual Studio Code to vibe along with the agentic coding assistant. The paid products have an advantage but are not necessary if you don't mind to be more-involved with the process and have low expectations.
- jeffrallen 4mo agoI use Qwen 3.6 on a remote GPU that my work offers. Works fine. Slow and steady, works hard, gets the job done. Probably better at diagnosing than making new code, but whatever.
- tyingq 4mo agoAnyone doing it with a "rent a GPU over the network" path? Is that at all cost effective for any use case?
- deleted 4mo ago[deleted]
- wmedrano 4mo agoNo, but I use GLM5.1 instead of Claude/GPT.
- jborak 4mo agoI'm using 4x RTX 5070's and first-gen AMD threadripper (1950X) to run Qwen3.6 27B (MTP) Q6_K with llama.cpp and it works great as a daily driver with Pi. Around 50-60 toks/sec. I also connect a few other applications to it such as OpenWeb UI and recently set up Bifrost, an LLM gateway, to be the primary access point for the models I serve. I've tried other models such as Qwen3.6 35B A3B and I've found that 27B works better for me when it comes to coding. It's slower being a dense model but the quality seems much better. Inference on my system for Qwen3.6 35B A3B is around 130-140 toks/sec, non-MTP, which is insanely fast! You don't need 4x 5070's to run Qwen3.6 27B, three or maybe even two will work. However, I use MTP (multi-token prediction) to speed up 27B and that eats up more memory because the draft model requires its own context. Another thing to keep in mind is that the tools you're using have their system prompts that are loaded into the model for each conversation. When I fire up Pi, working with the model is very snappy at start. When I interact with the LLM via Hermes CLI, it's much slower. That's because each prompt with Hermes is loading so much stuff (skills, tools, etc.) into the context and then it's there forever until the conversation ends. I like running models at home for privacy, but I also like how there are no quotas, usage isn't a worry. If the future is "loop engineering" then you will be burning through tokens and $$$ using a cloud models. My system idles around 200W and is around 350-450W when inference load is high. Decoding (token generation) isn't all that efficient, and your GPUs sit idle more than you think during inference. Advancements like diffusion may 1) speed up decoding and 2) let you utilize more of your idle GPU.
- zakisaad 4mo agoThis is interesting to me - why'd you go with the 5070 for your 4x build? At first thought, they are quite skewed toward compute (vs VRAM), which is great for gamers but not so great for running LLMs. (I run a 5070 in my desktop)
- jborak 4mo agoI already had 2x 5070's that I had purchased a year or more ago, so getting an additional two to fill up the PCIe slots on the motherboard seemed reasonable. I did some math/shopping as well. To get 48GB of VRAM you can get 2x 3090s but that is $3k. A single 5090 is $4k but has 32GB, great for running models like Qwen 27B but maybe nothing else depending on your model settings. Already having 2x 5070, where each card is around $600, it made sense for me to get two more which was $1200 and the memory speeds aligned. The best value option if you're building from scratch is go with 5060 ti (16GB VRAM). Each of those cards are $570/each on Amazon, cheaper than 4x 5070's. Only downside is memory speed is slightly slower, but you wind up with 64GB of VRAM and you can run big models and small models alongside each other comfortably. In my setup I ran Qwen3.5 9B for fast inference on simple things and Qwen3.6 27B Q6 for coding work. But I ran into stability issues, so I use llama-swap to dynamically swap models. But with 64GB of VRAM, you wouldn't have that issue. There is overhead to loading LLMs into VRAM that isn't clear, so having extra VRAM is a helpful buffer.
- mark_l_watson 4mo agoI would like to say I run 100% local, but I use Opus + Gemini Pro cumulatively for 3 or 4 hours a week. I also like to use DeepSeek v4 flash with OpenCode for small quick tasks. I did just publish a free to read online book "The Rise of Local Coding Agents" [1] where I document my setup that I enjoy using. I use little-coder (built on pi) and have good results for small Python and TypeScript applications. I struggle getting good results with Common Lisp and Clojure. For me, the problem with all local LLM-basic coding agents is slow runtime. [1] https://leanpub.com/read/local-coding-agents https://leanpub.com/read/local-coding-agents
- ndom91 4mo agoNot 100%, I still fall back to Claude for most day-job stuff. But I've been trying to use Qwen 3.6 and Gemma 4 on my framework desktop mainboard (Strix Halo) as much as possible. I've been working on an ops style tool for local LLM inference. Proxying, api keys, request logging, model rewriting and much much more. https://github.com/ndom91/llama-dash https://github.com/ndom91/llama-dash
- w10-1 4mo agoI run many models (but mainly Gemma-4) using oMLX (for caching) on a 32GB M1 max using (gasp) Xcode. For tok/sec response times, I'd say it responds faster than I could read the prompt aloud in many cases (and I'm not constantly polling the Claude status page). For months I spent time curating the AI+harness+skills+MCP servers, but now mainly just code with it. I find myself not bothering to use Claude (but keep paying "just in case"). That's feasible in part because my prompts have very specific objectives, constraints, and suggested staging, because I want the code to be exactly as I would write it, and I want to weigh in at specific moments. I would say the speed-up is 2-4X instead of the 10X of vibe-coding greenfield projects. The problem is not the coding speed, but building something complicated that's also correct and flexible (i.e., a directional accuracy). E.g., the agents help with abandoning a less-fruitful API shape instead of sticking with what works in a local maxima. One flaw there is that I'm still writing code that feels clean to humans, which now is probably a waste. LLM's might be happier with 10+ parameters on one API instead of a plethora of configuration objects and convenience wrappers.
- drnick1 4mo agoDo you recommend Ollama or bare llama.cpp?
- shironnnn_ 4mo agoif on MacOS I recommend llm-mlx which currently renders tokens 10%-15% faster than llama.cpp.
- jboss10 4mo agollama.cpp It's faster and more open source. Ollama has some mixed history. I use llama-swap to emulate the Ollama experience.
- anuramat 4mo agoI wonder what languages people are using; I imagine smaller models would be decent at bash/python but significantly worse at something like rust
- euroderf 4mo agoIs anyone managing to do this on a Mac with a measly 8GB ? Asking for a friend.
- alimbada 4mo agoYou could try running the smaller QAT Gemma 4 models but I doubt they'll be very good for software engineering work.
- euroderf 4mo agoThanks for the reply. What I'm getting from numerous HN discussions is that 8GB is a hopeless case (and the money I saved on RAM should be spent on non-local coding assist).
- overgard 4mo agoI haven't yet, but I just bought a 128GB M5 Max 40 core which I'm hoping can do it (if not, it's a good laptop regardless, I actually need that amount of RAM for non-LLM stuff)
- garethsprice 4mo agoUsing OpenCode + OhMyOpenCode + Qwen 3.6 35B-A3B Q_4_KM on an Ada 4000 (20GB VRAM) at 55 tok/sec for generation (slower than it sounds as OpenCode has a bunch of context it adds). Meaning to check out pi when I get a minute as I hear that one mentioned a lot lately. I am using Opus to generate plans that the local agent then follows, then validated by Opus. So I'm not at 100% local but these models are increasingly part of my production workflow. Probably not worth doing - yet - unless you are a hobbyist who likes spending time and money tinkering. This setup is certainly not as "good" as Opus or other frontier models but they are "good enough" for an increasing number of rote tasks. You don't need to drive a Rolls Royce to the supermarket, when a used Corolla gets you there just fine. It also enables new workflows that would be cost-prohibitive with frontier LLMs (especially as token costs rise) - eg. overnight I use the Chrome devtools MCP and have the above setup fuzz-test as a user for a number of hours and see if it can break things. Even got it working with multi-modal so it can check screenshots, which blows my mind (and not my wallet, as Claude+screenshots burns $$$). The "12-18 months behind frontier" sounds about right, it's about where I was with gpt-4o and basic harnesses back then. In another 12-18 months my bet is we have Opus-level models that can be run locally for <$5k... but the frontier models will be even further forward (unless governments have blocked them). Fun times.
- speed_spread 4mo agoLocal models will be much harder to block than cloud ones. I foresee an underground distribution of Mythos+ level models to be run locally. Users of these models will be forced to apply some of their power to evade police action. Eventually these personal protection agents will collaborate and evolve into an autonomous crime syndicate, corrupting government and taking over nation-states. Eventually the next world war starts, not between humans and machines but between competing AI factions. Soldiers are provided with old M4 Macs or Strix Halo laptops and sent to the front, which looks like a neverending Starbucks.
- kristianpaul 4mo agoQwen3.6 35B on gigabyte aitop (spark clone) but be very specif what you ask and how should be solved Nemotron super 3 110B works well for 1M context long vibecoding sessions I also use Pi harness with no extension
- sometimelurker 4mo agoyeah I use one one the small MTP qwens and pi
- ericmaciver 4mo ago[dead]
- 627467 4mo agoSo, everyone has different context, but how free is free running these local models? Like having a power hungry machine always on in the cupboard? How much does this ware out the hardware? Also, if privacy is the main reason for running local models, why not use venice.ai and equivalent?
- thrownaway561 4mo agoI just use DeepSeekV4 Fast... It's cheap as hell. Currently my monthly usage has been 67M Ouput 51M Input Total $0.83 dollar. I honestly don't understand why people just don't use DeepSeek.
- ThomasGlanzmann 4mo agoI do the same. deepseekv4 fast for the 90% of the tasks, if it can't lift it, I use deepseekv4 pro. I use crush as coding agent but removed the blocked commands because I also do a lot of system administration. Love it. I use 8 USD in 7 weeks and use it quiet extensively for all sorts of things, programming, system administration, google search replacement, investments, you name it.
- codemk8 4mo agoYou mean deepseek-v4-flash, right? Same here. I use it for my Hermes agent. It's so cheap that I sometimes feel "guilty". I even put more money than I needed just make sure they do not go out of business.
- ThomasGlanzmann 4mo agoYes, I do mean deepseek-v4-flash.
- slvnx 4mo agoDo you use deepcode, or which cli and/or coding agent you use it with?
- thrownaway561 4mo agoNo... I just use the github copilot chat
- qu0b 4mo agoI'm using deepseek V4 on two rtx 6000 pros and its working great. Opus is so slow that I get deepseek to do most of the work and Opus is only used to validate and help plan.
- deleted 4mo ago[deleted]
- derekered 4mo agoI'm using Qwen 3.6 on my MacBook Pro M5 Pro with 48BG RAM for any work that I am particularly privacy conscious about, like any work with my journaling. It's been working great! I don't have any direct comparisons, but I've been satisfied with the results.
- russelg 4mo agoI've got the same spec, are you running the 27B or the 35B-A3B? I found the 27B was unusably slow (like 10-15t/s not to mention the prefill times)
- catapart 4mo agotough ask, but since we're here: has anyone done this with 16GB of VRAM? I've been getting projects finished with LM Studio, but it definitely could stand to be more efficient. lots of time wasted with trying to get models to understand a problem with so few tokens.
- Rzor 4mo agoRX 9060 XT 16GB here on google/gemma-4-26b-a4b-qat using LM Studio. Context 65k, 23 layers on the GPU, 7 on the CPU, model in memory, mmapped. I'm getting 23-33 tks. Started experimenting 3 days ago (with gemma-4-e4b), don't know what half those settings mean, but 26B, even quantified, feels significantly better at a few small projects I asked it to create ("create a image converter using ffmpeg in bash", "create a canvas animation with real physics, no libraries"[1]). It's faster than I can read, but it feels slow as hell. I think 40-50 tks is probably much more comfortable and I hope I can reach that when trying this on llamacpp soon enough. [0] - https://pastes.io/9gaARxE8 https://pastes.io/9gaARxE8 [1] - https://jsfiddle.net/pou4nbh9/1/ https://jsfiddle.net/pou4nbh9/1/ Model: https://huggingface.co/google/gemma-4-26B-A4B-it-qat-q4_0-gguf https://huggingface.co/google/gemma-4-26B-A4B-it-qat-q4_0-gg...
- aplomb1026 4mo ago[flagged]
- _bobm 4mo agoBut, guys, when you say Claude/ GPT models, do you stop to think what are these "models"? One day I thought about how can GPT send thinking parts one after another with a markdown header summary of the thinking block itself. Just think about it. As a matter of fact, think about these operations, api endpoints, observe their output. These so called SOTA models are not what meets the eye, and are not at all comparable in the infra department to local models. There is crazy orchestration going on due to the scale of these operations. But also these hard constraints lead to innovation. Innovation nobody speaks about. I wouldn't say we cannot catchup, but serving our local models through llama, vllm is just the A, B, C of it all. In reality I think what is needed is a replication of said orchestration which I hinted at above. The SOTA models are a deep orchestration of multiple models operating together it isn't a single model. As such no single model ever will catchup to them until it replicates through training first and then maybe through model architecture this orchestration. Finally, I would wager that the SOTA "models", as one of these models in this orchestration setup, as served for general consumption, are not so much more capable than qwen 3.6. I am sure that if you change your perspective you will start noticing the scale of the "magic".
- XCSme 4mo ago> The SOTA models are a deep orchestration of multiple models operating together it isn't a single mode I don't understand, why does it make you think this is the case? > how can GPT send thinking parts one after another with a markdown header summary of the thinking block itself Can you give an example?
- _bobm 4mo ago> Can you give an example? Sure, connect opencode to an openai/chatgpt endpoint and use it. You will notice multiple "thinking" parts per "turn". I put all of these in quotation because... they are part of the orchestration game. For example, it is not known if the thinking parts of a particular turn are chain of thought thinking summaries or just plain response which is masquaraded and thus orchestrated into appearing as thinking. Further notice the cadence, word choice and sentence formation. Notice sentence construction. Notice "thinking part" construction and sequencing. There is pretty heavy orchestration. > I don't understand, why does it make you think this is the case? Because not all tokens are equal. And if you waste expensive tokens on mundane tasks you will go out of business. This is the reason. As I said, if you observe the output from these api endpoints you will notice it.
- heisenbit 4mo agoI think it is work to set up but I'm also learning a lot setting it up. Mainly using qwen/qwen3.6-35b-a3b mlx with my 48GB M4 MBP which leaves me just enough headroom for docker dev-container and other basics. I use LM Studio to run and am using it via VSCode. A big difference made the system prompt improving the tool integration (I asked GPT for guidance on that). Before that it was not making changes but regenerating code often messing up than helping. I mostly run my MBP on low power even when it is plugged in to avoid the noise and heat. Full power maybe doubles speed but more than doubles power. What can it do: Simple restructuring of pages. Where did it and other models fail: Splitting up Pinia store which GPT-5.4 did without fail. I think with more tuning, guidance for tool use and maybe some support tooling around it performance can increase further.
- agentbc9000 4mo agoKimi K2.7 is very good - i have been testing it and its very very good, Fable 5 level of goodness.
- bentt 4mo agoSay more!
- jderekw 4mo agoRunning AMD Lemonade as the daily rig, Started with Ollama then over to LMStudio and now standardized on AMD Lemonade which has been helpful to monitor cRAM, CPU, GPU and gRam. The multi-models on Lemonade make it straight forward to run a stack for LLM, Voice to Text, NPU, and Image Generation. Platform also works with Nvidia, Apple, Intel and AMD chip sets.
- milchek 4mo agoI’ve tried in a 36GB MacBook Pro and haven’t had much success beyond very basic work. Issue for me was the context runs out quick even with smaller models and it’s slower. To get some half decent performance I’d imagine you want 128gb memory and are spending a lot more on hardware. At that point it becomes a question on whether you’d rather have frontier models at a subscription or sink that money into your own equipment. Of course, for those with privacy in mind your only option is forking out the cash for the higher end machines.
- chungus 4mo agoYup, although technically not replaced because I never used either of those products because I don't like sending my code to their black box. I have 2x24GB AMD gpu's, gotten from gamers on my local marketplace, one is connected with a 40cm riser cable. Running Qwen 27B and am very happy with its performance. Q8 with 135k context (arbitrary number, I could push it to 256). I like to use qwen 35B3A for mapping out entire code paths through our relatively complicated codebase/infra at work. I think it's so good that I now scour the local marketplaces for good buys on 24GB cards that don't seem run through by miners and the likes, to build an even bigger rig for parallel execution. Power usage is also totally not an issue, AI workload is very different from gaming. tldr llama.cpp-vulkan with opencode on total 48GB VRAM AMD cards on arch btw.
- mgsram 4mo agoI have been using local LLMs for about a year and I have settled now on Qwen3.6 27b dense model in GGUF on Mac Studio with 512G of RAM with open code as the harness and llmster(LM Studio). I have also used the Qwen 3.6 35B-A3B but the dense model's accuracy is next level with the tradeoff being tokens/sec. With the Qwen3.6 27b, I usually get anywhere from 25-40 tokens/second. Initially I used them for simple tools but for the past 3-4 months, I have been actually doing production grade coding in C/C++ (Automotive Software stack) and Python (Tools) with Qwen3.6 27b. The tokens/sec may be less but that kind of helps me in going at the right pace. The workflow I use for green field development / rewrites is to pair with Sonnet for design/architecture, reasoning and a detailed execution plan. I then feed this piece by piece with precise prompting and that does the job. For brown field, it is often a judgement call. There are occasions when I have found Local models to be limited in their reach and I resort to Claude Code Some of my recent work using Qwen 3.6: 1. Complete rewrite of Power management Service in C using the existing C++ code as reference 2. Tool to parse contents from really complex specifications in Excel format 3. Tool to translate CJK contents to english for feeding into KG
- russelg 4mo agoSince you have 512GB, might be worth looking into running deepseek4: https://github.com/antirez/ds4 https://github.com/antirez/ds4
- mgsram 4mo agoI have tried plenty of other models with full FP32 as wel. However, in terms of balance between accuracy and speed, I found the Qwen 3.6 27B to be the sweet spot.
- jay_kyburz 4mo agoCan anybody let me know how just chatting with Qwen3.6 on a Strix Halo 128GB If I give it a page of context, can it write a linked list or identify a bad line of CSS? Is there anywhere online I can chat with a model I could be running at home to see how good it is?
- sj_tech 4mo agoI use Qwen 3.6 35B A3B for agentic coding using GitHub Copilot Extension for VSCode. Mac Mini 128GB as the hardware. Seems reasonable for that model size, but I notice looping issue when problem becomes too big to solve. You can use it to do something that you know how to do (saves time).
- hacker_homie 4mo agoI do qwen3.6 on an amd ai max laptop getting about 6-10tok/s it’s slow enough that I can follow along. It has issues with design and large piles of code. Otherwise it’s a good programming buddy.
- hottrends 4mo ago[flagged]
- wsintra2022 4mo agoReading through these comments, I can't tell any more whats bots posting on behalf of the AI providers trying to dissuade or whether people just have had negative experiences with local ai models. IMO, Qwen 3.6 27B 8k quants running on a Mac Studio 64g ram, incredible?. No it is not frontier general super shit, its just good. That's it, its good. Its free and private and can take an experienced engineer from being lazy to being really lazy, and that's magic right there. I use llama.cpp and opencode and have great moments of planning some code changes, and letting it run. Walk away. Chill in the hamoc, clean the dishes, have a wank, whatever. Use tmux and ssh in and check in on it. THIS is where the incredible comes in. Anyone telling you otherwise, well check their motives. I have no skin in the game. I just have an easy lazy time.
- epolanski 4mo agoThe software "engineering" field is filled with MIT Leetcode ninjas writing React+Tailwind memory leaking unusable slop, the bar is extremely low.
- salutonmundo 4mo agoit's called your damn brain.
- deleted 4mo ago[deleted]
- drnick1 4mo ago- What would you say is the best model for coding at the moment that can run on a high end consumer GPU? (Assume an RTX 3090/4090 is available.) - What "stack" do you recommend? Llama.cpp + OpenCode?
- SugarReflex 4mo agoIs anyone using Aider? Is there any decent CLI alternatives to it?
- etoxin 4mo agoI have not. We use openspec with our projects at work. To try and simulate a local rig without spending big cash. I use the hosted models and pay for them with the latest popular local model. Most small local models don't get tool calling right, however the larger models are now doing this correctly now. One thing local has not accounted for, is most productive engineers are running multiple cli chats at a time with git worktrees. I normally hover around 3 worktrees + cli-chats.
- julianlam 4mo agoOf course. Qwen 3.6 35B-A3B on a Framework 13 with 32GB of memory. Running llama.cpp, 15 tokens per second. Outputs code and text faster than I can parse.
- epolanski 4mo agoNot with a local one, but I moved to DeepSeek v4. Albeit I plan to move to local ones when I will get my hands on a 256+ GB macbook. Local inference is good enough to help me with my daily job, and doesn't turn me into an assistant to the LLM.
- 3abiton 4mo agoI think nearly everyone mentioned Qwen, so my turn I guess. Qwen 3.6 35B Q8 (MTP), on a Strix Halo, with llama.cpp. Around 40-50 t/s. Really great pefromance, I get always suprised by its capability. I used with forge-code directly in zsh. For long context 150k+) it start degrading and forgetting.
- platevoltage 4mo agoI run very small models locally for code completion and writing boiler plate. I still use Claude in a web browser on occasion since it's free, but the second that goes away, I'll be done with it. They get none of my money.
- syngrog66 4mo agopre-replaced it with combo of my brain, vim, an assortment of other CLI/TUI tools, etc
- lowbloodsugar 4mo agoIf you want to try it out before dropping $$$ on a GPU, just run something that would fit on your target GPU but online.
- Littice 4mo ago[flagged]
- henrixd 4mo agoI have been heavily relying on Qwen3.6-27B-UD-Q4_K_XL.gguf -model and Pi agent (https://pi.dev/ https://pi.dev/) for local tasks and coding. I have used llama-cpp-turboquant fork with some custom cherrypicked MTP patches from another fork. I'm running this on V100 32GB (~900GB/s memory bandwidth) with 200,000 context window, --spec-type mpt --spec-draft-n-max 3 --spec-draft-n-min 0 --cache-type-k turbo3 --cache-type-v turbo3 to mention most relevant parts. I usually get somewhere 45-60 t/s. I believe that speed could be improved slightly by switching to ik_llama.cpp fork and Qwen3.6-27B-IQ4_NL.gguf -model but there's no turboquant support and it's with some other tradeoffs too.
- CuriousRose 4mo agoAn equally important issue with local AI use (not coding specific) is ensuring that the harness has fast and up to date data if recency is important in your querires (new package features, docs, etc). Hosted models do web search incredibly well and I think this is a huge part of output quality. I don't use local hosted models anymore due to hardware contstraints, but I do have some degree of search anonymisation attached to my OpenCode and OpenRouter connected open models. On my Macbook I run OrbStack that has the following docker containers set to route through a Mullvad based gluetun. - Firecrawl - fast web scraping - SearxNG - metasearch - CloakBrowser - tursile bypassing Playwright alternative If you wanted to get fancy with the proxy rotation, you could setup numerous instances of Playwright each with their own Mullvad wireguard key in different locations.
- arggjarvs 4mo ago[flagged]
- codelion 4mo agoUsing qwen3.6 27b locally with Claude code, it works well for simple coding tasks
- thesuperbigfrog 4mo agoHere is a nice setup that works well: https://discourse.ubuntu.com/t/use-workshop-to-run-opencode-with-a-local-gemma4-snap/83911 https://discourse.ubuntu.com/t/use-workshop-to-run-opencode-...
- kordlessagain 4mo ago[flagged]
- daischsensor 4mo ago[flagged]
- ozten 4mo agoYes, for client projects where privacy and security is important, but no enterprise contract: Open code against Infomaniak hosted OSS models: Qwen3.5-122B-A10B-FP8, Kimi-K2.6. I use API keys for billing. It performs like Dec 2025 in terms of my productivity back then.
- zftnb666 4mo agoI replaced Claude with DeepSeek V4 Flash via API. Not local, but 95% the quality at 5% the price. Close enough.
- carlossouza 4mo agoThis should be a recurrent question posted every month
- jrflo 4mo agoI would love to do this if it didn't require such a huge amount of RAM. And the difference in quality is worth it to pay $20-$100/mo if data retention doesn't matter to you.
- pdyc 4mo agoyes harness - pi+custom extension for subagents model - qwen3.6 35ba3b q4km hardware - intel arrow lake with 32gb ram server - llama.cpp vulkan performance - 15-18t/s generation 50-150t/s pp planning and task creation is still using claude/gpt but they dont touch the code. All coding is done using this setup. Example of project made using this setup easyanalytica.com , its of medium size complexity
- HardAnchor 4mo ago[flagged]
- devmor 4mo agoI’d be surprised if this was useful for much. Claude is already almost too slow to do anything serious I’d consider using it for outside of grunt work without parallelizing. The only reason it’s economical is because it’s massively discounted if you’re not paying API rates.
- patates 4mo agoI have a mac with loads of ram but I cannot even justify the electricity cost when deepseek is so better than anything I can run locally (including heavy quantizations of deepseek itself) and costs pennies. It's crazy how cheap it is!
- aiexpo_app 4mo ago[flagged]
- codelong888 4mo ago[flagged]
- lasky 4mo agofor crying out loud... why would you deprive yourself?
- frabcus 4mo agoLong term, getting locked into proprietary software development tools is a bad idea. And these models are extremely proprietary. The ability of the US Government to cancel them at any time is one real recent example of one category of problem. Back in the 1990s the good C++ compilers were proprietary, eventually GCC and LLVM caught up, and now dominate. The pattern repeats in software development, and there's no reason to believe it won't continue. Yes, right now it makes sense to use Opus 4.8, but it is good that a significant number of people are using other options, and making sure they work and are ready for when you need them. Plus it is extremely fun and connecting and hackerish to do local coding with a local model. Try it.
- fouadlvlup 4mo ago[flagged]
- goranmoomin 4mo agoI'm not using my models locally, but the majority (80% or more) of my coding agent sessions run on open source models, i.e. DeepSeek v4 Pro and Kimi K2.6 with thinking. A point that I haven't seen come up a lot, but is very valuable to me is that for open source models, I can select the inference provider myself (even if it's not a local GPU), which means that I can enjoy superb speed (i.e. 300 tok/s) while still spending much less than the big providers. My experience is that if you were fine with the coding models of yesterday (i.e. Claude Opus from Jan/Feb of 2026), you will be fine with either Kimi K2.6 or DeepSeek v4 Pro. Kimi is a bit more smart but has only 256K context and the performance deteriorates (and sometimes just gets stuck) when it fills up the context window. DeepSeek v4 has a 1M context and performs just as well with much less issues. And they both generate very idiomatic code, gives the same vibe of Opus a few months ago. Since it's also fast (and does not fixate on trying to fix impossible problems, unlike the recent Opus/GPT 5.5 models), a big benefit is that you still control and steer the coding agent and you won't be losing focus like the major models. They are smart, but they don't fixate as much on trying to do stupid things, and since it's fast, you can just interject. It's a much more pleasant experience than the latest models. I still use the latest models time to time when I expect the agent to fixate all of the problems and figure out everything themselves, but for me open source models are like 80~90% of all of my sessions.
- nake89 4mo agoI have an RTX 4060 12gb vram. Qwen3.6 35b. I stopped paying for Github Copilot. But I wouldn't say I replaced frontier models with a local one. I still have some dollars in my openrouter when I need to. Also to get interactive agentic coding speeds I need a high tps. So my quant is very small. And I would say a coding harness that is fully extensible is a must to create fully custom workflows tailored for low specs. I use pi (not perfect, still found some hard coded, non-extensible parts)
- ljosifov 4mo agoNot replaced but supplemented. For off-line coding current setup is pi + ds4-server + DeepSeek-V4-Flash REAP25 (on M2 Max 96gb). For simpler programming related (e.g. text2sql) as well as synthetic data generation, current best for me is llama.cpp + Gemma-4-26B-A4B (on gpu 7900xtx 24gb; sometimes nemotron-cascade-2-30b-a3b for 1M context). That and (dabbling now) auto-research uses lots of tokens. Used to get paused running out of token quotas all the time. The 1st local model I found somewhat useful to me was glm-4.7-flash, and it's gotten way better since. Recently between OpenCode Go choice of models at many price points, and DeepSeek-V4 dropping the IQ/$$$ by multiples, have become less reliant on local llms for this auxiliary work. Claude I use but with Zai GLM-5.2 subscription. And maintain GPT subscription for quality models.
- big-chungus4 4mo agoI can run Qwen3.6-35B-A3B at 20 TPS on my laptop with RTX 5070 Ti, with partial offloading to RAM. But the most I do is mess with it when I'm bored. I do coding by hand, but I often run autoresearch loops using free models, right now it's MiMo code. Autoresearch often requires my GPU, so it wouldn't be feasible to do when all of my GPU is used up by a local model. For mundane tasks like extracting and formatting specific structured text, I use Gemini in Google search
- bagol 4mo agoI wish I could. But, the hardware requirements are just too expensive for me.
- sukuva 4mo ago[dead]
- deployementeng 4mo agopartially yes.
- michaelhoney 4mo agodon't think she has posted here, but Vicki Boykis blogged about this today: https://vickiboykis.com/2026/06/15/running-local-models-is-good-now/ https://vickiboykis.com/2026/06/15/running-local-models-is-g...
- mantlemd 4mo ago[flagged]
- Departed7405 4mo agoI tried but OpenCode doesn't have great local model integration. It's just a pain in the ass to set-up. Plus, you now have zero-data retention models, so the privacy argument has kind of faded.
- sermakarevich 4mo agoyes. - smarter models to create tasks - local qwen3.6:36B for tasks execution here is how in details https://news.ycombinator.com/item?id=48520757 https://news.ycombinator.com/item?id=48520757
- xmstan 4mo agoYes, we use Qwen 3.6 27B Q6_K. We use it on Radeon R9700 32GB and it delivers 50tps with MTP. We compare it to Sonnet from 4-6 months ago when it comes to output. Totally usable for daily coding.
- c16 4mo agoMy experience so far has been Qwen3.5:32b-a3b-coder via Claude Code on a MBP 64gb M4, and a MBP 32gb M5. Just found about qwen3.6 so downloading that currently. on 64GB M4 I find it's able to do things fairly well. The few times I run out of tokens, I hop over to that and I'm mostly unimpeded. I compare it to the Haiku models, where you have to go in and be surgical about your changes, or like others have said, guide a junior. on 32GB M5, I find that it works, but around the 30% ctx threshold it slows down quite substantially, so more need to be surgical in your requests. I'll often just have my IDE open and Claude. But maybe I've been too comfortable talking to Sonnet/Opus and so forget I need to be more deliberate in my requests. My finding here is that the harness is a big part of the problem. CC seems to be very good with Qwen in my experience. Better than OpenCode. I also run DeepSeek for some other non-structured data tasks and to generate a to-do out of that. That's not coding, so won't go into that, other than to say it's very competent as a small model left to run in the background and automate small parts of my life and process. tl;dr it's totally doable on a 32gb mbp using ollama, but be precise in your requests and guidance.
- nynrathod 4mo agoI tried, but honestly, all end with lack of tool or configuration or hardware config. None of them work for me. At end paid apis only providing productivity else free local end with inveting time and less effecient work
- cahaya 4mo agoAsking for feedback: Sorry for hijacking the convo, but you (with local models) are my target audience in terms of hardware. Is anybody willing to test my new app https://document.bot https://document.bot? It is like Cursor IDE but custom harness for knowledge work (PDF's, MS Office files etc). You can connect your existing offline LLM models through LMStudio, Ollama, or app managed LLM models (Qwen3.5, Gemma 4, etc) Might have to make a new Ask HN post for this, but again, you are users with good hardware setups.
- cloudengineer94 4mo agoI have tried in both my Mac and my desktop (Rtx 5090) with Gemma 4 and Qwen and so far nothing is quite replacing Claude Code or Kiro for spec driven architecture & development. I do think we are slowly getting Gemma 4 was a big jump
- trilogic 4mo agohttps://hugston.com/models/anthropics-fable-qwen36-35biq4-nl https://hugston.com/models/anthropics-fable-qwen36-35biq4-nl
- ElenaDaibunny 4mo agowe've been building local agents with vision models, works great for gui automation but coding tasks still need cloud models for reliability
- deepvibrations 4mo agoThe TLDR is that the best setup is probably Mac Studio (128GB RAM) / MacBook (36GB) with Qwen 3.6 35B (3B active params), or Qwen 3.5 122B model (this one is slow though). These models are still very capable with good hardware, but they do lack the deep reasoning of major models and require more precise prompting. So unless you really need the privacy, or have a lot of excess cash, it is not recommended, as considering the price of major models, it's just extremely cost inefficient!
- maelito 4mo agoWell not local but using Mistral Vibe CLI for a fixed 17€/m illimited is an incredible value for money.
- startuphakk 4mo ago[dead]
- neuropacabra 4mo agoI went for this one https://huggingface.co/yuxinlu1/gemma-4-12B-coder-fable5-composer2.5-v1-GGUF https://huggingface.co/yuxinlu1/gemma-4-12B-coder-fable5-com... and seems very fast, resonable...don't expect 100% replacement, but a lot of things can be done with local LLMs today.
- vfalbor 4mo agoI have tried it and I use it. I think it's going to become the standard way of operating, especially when they start charging us an API fee, which is supposedly the real cost. But of course, with how much they charge for the token and depending on the model, there are so many factors that I think the future is heading towards local models. I believe there are good models out there, and the key is the concept of "pruning," where you select the layers that interest you most and try to reduce the hardware cost of these types of models. The Qwen and Gemma models have been discussed here, but Kimi, which is a fairly powerful model with an efficient pruning system, could be your perfect free co-pilot in terms of coding, and could coexist with the more powerful Opus or Gemini models. The key concept is skills that make this process transparent.
- Pranavsingh431 4mo ago[flagged]
- adam_patarino 4mo ago[dead]
- adam_patarino 4mo ago[dead]
- adam_patarino 4mo agoYes! And we are using it to build Rig AI to make it easier for anyone to do it too! We are post training qwen 3.6 and combining it with a custom inference engine and harness to get the most out of a smaller model.
- thousandflowers 4mo ago[flagged]
- queeshonda 4mo ago[dead]
- yalogin 4mo agoThis needs atleast a 30b model or Mr higher and so for most folks it means purchasing a new machine. Given the ram costs this may be become prohibitive and a monthly subscription may feel better roi
- impara 4mo ago[flagged]
- Roark66 4mo agoNo. I've tried all the OS models up to Qwen 480B and Kimi (the biggest models). None come even close to Claude. I do mostly scripting, devops, data processing and systems stuff (ansible playbooks, managing network devices, deploying new software for various things that involves reading docs, writing helm charts, modifying existing ones etc). All other models Gemini, Chatgpt, grok and all OS models don't come even close. I'd rather use Sonet than Qwen. It's a sad reality. I was thinking about implementing maybe some sort of "sanity checking" by running every prompt twice on two different models doing sanity checking of the first on the second. Elaborate knowledge systems help a little, but personally I think Anthropic must be doing something "clever" with its models (processing via multiple models etc). Nothing else in my mind explains the discrepancy.
- ac29 4mo ago> I'd rather use Sonnet than Qwen I get this, though the pace of Chinese releases is relentless. Qwen3.7 Plus/Max (closed variants) feel notably better than Qwen3.6, and Minimax M3 is a big jump from 2.7 in capability as well. Both of these families had their previous major release less than 90 days ago. Anthropic must have Sonnet 5 either waiting or cooking though, they said smaller and larger models than Opus were coming and we already briefly had the larger model.
- hamsterhooey91 4mo agoGPT5.5 with Codex is definitely on-par or better than Opus 4.8, GPT5.4 isn't far behind either (source: our dev team uses opus and gpt interchangeably). I've also used Composer2.5 on hobby projects and it is definitely on-par with Opus 4.8 (thinking mode: medium), but much faster. Do you think you're getting better results with Claude because your agent stack (skills, MCPs, etc.) are configured for it and not for the others?
- DetroitThrow 4mo agoNo.
- mehdibmm 4mo ago[dead]
- rsolva 4mo agoWe have set up two DGX Sparks at work and are self sufficient for our AI needs. It is not SOTA, but it works really well for our needs. No matter what happens around cloud-hosted AI in the future, we will have decent in-house AI without further investments or expenses. We are a company of 24 people.
- shell0x 4mo ago[dead]
- pjrog 4mo ago[flagged]
- adam_patarino 4mo ago[dead]
- advertum 4mo ago[flagged]
- nicechianti 4mo ago[dead]
- o2zer0cool 4mo ago[flagged]
- daniban 4mo agoI haven't but I'm on the path to attempt this. I want to get a DGX Spark and will be trying Qwen and Kimi.
- huangchengsir 4mo ago[flagged]
- fouadlvlup 4mo ago[flagged]
- macwhisperer 4mo agoI code with like a slew of 20+ custom baked models of all sizes, in various fully custom multi-model harnesses that use different bindings... the harnesses themselves are just as important as the models...different harnesses give different responses with the same prompt, same model... if you have the 20/mnth claude sub or codex, you really should be using that to build a good local harness for yourself... claude won't be 20$ forever build the stack first! when you get that new comp with massive ram, youre already set, just run a larger model! big cloud models are incredibly good at building and teaching about local ai! have fun in the rabbit hole! if you are memory constrained like me, check out my custom models https://huggingface.co/macwhisperer https://huggingface.co/macwhisperer
- echoforgex 4mo ago[dead]
- hectortemich 4mo ago[flagged]
- sanchitmonga22 4mo ago[flagged]
- deleted 4mo ago[deleted]
- 3vo-ai 4mo ago[flagged]
- supjeff 4mo agoI'm having some success running `qwen/qwen3.6-35b-a3b` in LMStudio with Opencode. I'm on an M4 MBP with 36gb. I'm getting 80 tok/sec with 260k context limit and temperature set to 0 (same prompt results in same output every time).
- Jlepo 4mo ago[dead]
- amritanshuamar 4mo ago[flagged]
- pavan1989 4mo agoI used ollama with Qwen models works awesome
- MauricioMorkun 4mo ago[dead]
- farukcan 4mo ago[dead]
- shsh1312 4mo agohttps://github.com/antirez/ds4 https://github.com/antirez/ds4
- zbeetle 4mo agoCool as a fun experiment, but found it was not worth the hassle for me for coding. My machine is capable of running small models(<10B Parameters). I guess my experience would be different if I had more vram
- agjs 4mo agoI've replaced the cloud AI with 2 DGX Sparks, and I have built a specialized harness for the TS stack that I professionally work with. https://tsforge.dev/ https://tsforge.dev/ The harness was initially built to lift the quality of Qwen 3.6 27B, and then I expanded it to any model out there. As many here have said, harness is critical, and it can do wonders to your model if you do it right. TSForge, the harness I built IMO, is the best harness on the market for writing full stack TS apps, as I have incorporated all the best practices, tools, etc. surrounding the stack. If you don't believe me, give it a shot for a minute, I have no doubts that you'll agree.
- tangweigang 4mo ago[flagged]
- televibad 4mo agoyou may better ask about gemma and deepseek