7 ms·
How to setup a local coding agent on macOS
- aplomb1026 4mo ago[flagged]
- cdolan 4mo agoIs there a link to the video? It did not render when I went to the page. Curious about the real-time feel of this
- dewey 4mo agoThat's the direct link: https://ikyle.me/blog/2026/how-to-setup-a-local-coding-agent-on-macos/Gemma_4_-_Short.mp4 https://ikyle.me/blog/2026/how-to-setup-a-local-coding-agent...
- c-hendricks 4mo agoNote this is cut to just before the model responds, so not a great way for people to judge the real-time feel of this.
- freerunnering 4mo agoThe full video is on Twitter: https://x.com/Freerunnering/status/2065275403548168398 https://x.com/Freerunnering/status/2065275403548168398 Plus a followup one where you see me type the question in and press enter (though that video is with Qwen 3.6, not Gemma 4) https://x.com/Freerunnering/status/2065354101878055038 https://x.com/Freerunnering/status/2065354101878055038
- deleted 4mo ago[deleted]
- c-hendricks 4mo agoNot sure you really need huggingface-cli to download anything if you're just using llama.cpp. You can pass `-hf ...` and it will download the models for you. Set `LLAMA_CACHE` to change where the downloads go: LLAMA_CACHE="models" ./llama-server \ -hf unsloth/gemma-4-31B-it-GGUF:UD-Q4_K_XL \ ...
- dofm 4mo agoYes. -hfd for the draft model.
- c-hendricks 4mo agoNice, was wondering if there was a flag for the draft as well. Not knocking huggingface-cli, just find it's much easier for people to try out this stuff when they can just mise use --global github:ggml-org/llama.cpp LLAMA_CACHE="models" llama-server \ -hf unsloth/gemma-4-26B-A4B-it-qat-GGUF:UD-Q4_K_XL \ --host 0.0.0.0 \ --port 11434 \ ...
- dofm 4mo ago—no-mmproj is also pretty useful if you're doing this just to try agentic coding and you're not processing images/voice. Stops it downloading the multimodal projector.
- ig0r0 4mo agoI wrote a similar post some time ago just used ollama and opencode https://blog.kulman.sk/running-local-llm-coding-server/ https://blog.kulman.sk/running-local-llm-coding-server/
- sleepybrett 4mo agoactually useful and the ollama gui could probably even simplify this more.
- takethebus 4mo agothis is the way, given anyone could swap for oh my pi / pi / etc
- mark_l_watson 4mo agoyes, whether for home experiments or at work, it is good practice (good hygiene) to be able to swap out both agentic harnesses and models. It is important to have a good strategy for exporting skills, etc.
- ig0r0 4mo agoyeah, I am using Pi right now, I switched from OpenCode
- carter2099 4mo agoI'm considering this right now. Is it very difficult to adapt to the new philosophy? Are you running mostly interactive or programmatic?
- ig0r0 4mo agoWhat do you mean by new philosophy? I use Pi the same way I used gpt or sonnet, just for simpler tasks.
- dofm 4mo agoUseful stuff in here that I wish I'd seen a few days ago :-) I am not convinced that the MTP setup for the QAT model adds very much in terms of speed on my M1 Max, but it is definitely worth experimenting with. Fiddling about with local models has done so much for my conceptual understanding of what is going on. FWIW and YMMV but I also found the Gemma 4 MTP head was occasionally breaking markup in Opencode, causing the thinking to display untidily and ultimately in some cases missing the stop token. So I've stopped using MTP there for now. Recent Qwen 3.6 models have developer role support so it will occasionally surprise you with a structured multiple choice questionnaire.
- mft_ 4mo agoI found a marginal downside to Qwen3.6-35B-A3B-MTP vs. the non-MTP equivalent on an M1 Max. I’ll maybe experiment with settings further though.
- dofm 4mo agoYeah. I think it might speed up time to first token but I am not sure how much that matters. I do enjoy their different personalities when they are tackling "explain this" type puzzles, though. Gemma writes so well — like a concise code blogger. It makes you understand that the thing we hate about AI slop writing is specifically the cheesy, marketingese sycophantic ChatGPT tone. It's a choice to sound that way. Qwen writes more tersely by default, like much english language documentation in Chinese open source projects. A couple of lines, code example, fact, code example, line of blurb. I use this prompt every now and then with a new model. It's obviously a classic SQL puzzle but I've asked new web developers this in the past (prompted by discovering that a client's subcontractor didn't understand it and was therefore unable to migrate some code from relying on dodgy pre-MySQL 5.x behaviours) — I have a MySQL 5 table like this: [id, label, category, score]. It contains a list of items in different categories (text names like cat1, cat2, cat3) with a numerical score. Is there a way I can write a SQL query to find the item in each category that has the highest score, without using a subquery? No two entries in any category share a score. — I enjoy seeing what it deduces from the subtext. Without "thinking" mode on, they always initially fail and you need to prompt them to find the answer. With thinking mode, they both produce really nice explanations. For me, as an old freelancer who is pretty cynical about vibe coding or "agentic engineering", what I really want is an AI tool that can help me start to solve problems and help me find the right terminology or generate some boilerplate I can tinker with. Both of these models do fine at the kind of "starter" writing that I want when I am trying to untangle an idea.
- namnnumbr 4mo agooMLX (https://github.com/jundot/omlx https://github.com/jundot/omlx) makes running the mlx inference server quite easy for those interested in UI-based hosting. oMLX also supports mtp or dflash drafting.
- w10-1 4mo agoAgreed (not sure what you mean by UI-based hosting). oMLX does the caching I need to fit models that are near gross memory, and it handles most of the work in finding usable models. After cobbling together various solutions over months, I now just use oMLX, often from Xcode. I can tell the difference between Gemma-4 (local/free) and Claude (paid) only on the largest tasks.
- amboo7 4mo agoWhay about of the tons of caches that just pile up until you notice that you must delete them manually?
- reddit_clone 4mo ago>64 GB Thats the rub. I have an M4 with 48G. I wonder if it is worth testing this out. My past attempts (with Ollama and various LLMs) were too slow to use.
- hkchad 4mo agoI have a M5 MAX with 128, local models are toys compared to hosted ones. I've spent a lot of time and money trying to make it work even 1/2 as well.
- iluvcommunism 4mo ago[dead]
- deleted 4mo ago[deleted]
- dofm 4mo agoIt all depends on what you want to do, I guess. If you're seeking the kind of hands-off claude experience, obviously not. They are slow. If you want to learn how these things work, train them locally, tinker, play with the code, grasp the fundamentals, or just out of sheer bloody-mindedness and principle refuse to tether the functioning of your application to a cloud API...
- jillesvangurp 4mo agoFrom an economical point of view, there's almost no point to using these locally running models. The only things they are good for would be dirt cheap using the smaller/older models via some API as well. Recovering the investment for the hundreds/thousands you spend extra on hardware easily funds a lot of that. Unless you are using this stuff at scale, it's probably not going to be worth it. I've dabbled with Qwen 3.x and Gemma 4 models a bit. They are alright but not that impressive. And my mac gets super hot if I use them for extended periods of time. It's just not very nice to use locally.
- leemoore 4mo agoI have the same processor and ram. The dense 30b ish Gemma/Qwen really don't break 10 TPS with or without MTP. MOE's in this range feel more usable if they are smart enough for your work. Probably would still use hosted versions of these over local unless. MOE's feel somewhere between sonnet 3.5 and 3.7 to me. Dense feels between sonnet 3.7 and 4 in basic coding or local agentic capabilities (not close to those in chat or world knowledge)
- attogram 4mo ago8b max on a std 16gb macbook. Anything more and your mac is toast
- benbojangles 4mo ago70b on my M1 max 64gb
- Obscurity4340 4mo agoHow much did tha thing cost?
- Aurornis 4mo ago> The benchmark prompt was: > Write a compact Python function that parses a unified diff and returns the changed file paths. Then explain two edge cases. > Each benchmark generated about 128 tokens. Generating 128 tokens is probably not enough for good benchmark results. MTP speedup depends on how often the predicted tokens are accepted. In my experience, the very early output has a higher acceptance rate, so short testing can give false positive speedups. llama.cpp includes a tool specifically for benchmarking that will sweep the arguments for you so you don't have to restart the server and send it prompts: https://github.com/ggml-org/llama.cpp/blob/master/tools/llama-bench/README.md https://github.com/ggml-org/llama.cpp/blob/master/tools/llam... EDIT: Also the section about downloading the models should have mentioned that llama.cpp has a "-hf" argument that will download the models for you. I appreciate the author for sharing their experience, but for beginners this might not be the best guide to use.
- liuliu 4mo agoRealistically, you need to experiment with any user prompt + a good amount of system prompt (at least > 1000 tokens, but realistically, in the range of 3000 tokens probably good). llama.cpp includes tools for that, what you are looking at is to have a prefill before token generation to measure it properly. Increasingly also, measuring token generation speed at longer context (32k or 64k) is important too.
- reactordev 4mo agoThis is akin to saying “it runs on my machine” without actually examining the problem. Sad. You’re absolutely right that 128 tokens is nothing, it’s a little more than a hello response.
- willXare 4mo ago[flagged]
- lloyd-christmas 4mo agoI thought the same thing when I started using locals, but the reality is that - for a given context depth - the token generation speed doesn't change whether it's 128 or 8000, it just lengthens the benchmark run time.
- vladgur 4mo agoI have used omlx.ai with great success to both download multiple mlx models (including gemma and qwen) suited for my hardware AND to be able to automagically launch both open-source and close-source (claude code, codex) harnesses using these models. All from a web or desktop UI You would not need to follow a blog post with omlx IMHO
- fridder 4mo agoIt truly is the SOTA for local inference on mac. Even when there are regressions the dev(s) are insanely responsive. It is the most impressive opensource project I've seen in a awhile
- benbojangles 4mo agoOmlx needs to incorporate macos native shortcuts use - macos can almost instantly extract text from pdfs and a bunch of other things using it's ane neural engine keeping unified ram for llm use. The two together would be awesome
- Dotnaught 4mo agoIn case anyone is looking for a sandbox to go with oMLX and Pi: https://github.com/Dotnaught/pi-sandbox https://github.com/Dotnaught/pi-sandbox
- dofm 4mo agoThis is useful. I'm still tinkering with Multipass VMs because I need the whole VM environment anyway and I'm on Sequoia. But I'd be interested if you did anything like that with Apple's container CLI instead; sooner or later I will have to upgrade to Tahoe because I want to play with the container CLI (and apfel).
- zmmmmm 4mo agoit looks handy but ... sbx policy set-default open just so the single pi sandbox can talk to localhost? ... this gives me some grave doubts about the rest of it being set up well.
- jmkni 4mo agoFYI you can open Claude code in the terminal, point it at this article and just tell it to "do it", if you're feeling extra lazy
- echelon 4mo agoThis is the way. I'm not Googling much of anything anymore. 9/10 times the information is awful, it's hard to parse out of whatever other spam it's surrounded by. Meanwhile, Claude will just do the thing one-shot or with a tiny bit of refinement. The gateway to knowledge and getting stuff done is the LLM. Google Search is a dinosaur. It feels like we're living a century into the future. Not even smartphones were this cool.
- tobyhinloopen 4mo agoClaude “respond in a friendly way that I agree with this comment”
- kingofthehill98 4mo agoYeah, if the future is "Claude, think for me" I'm happy to stay at the good old present.
- deleted 4mo ago[deleted]
- echelon 4mo agohttps://en.wikipedia.org/wiki/Is_Google_Making_Us_Stupid%3F https://en.wikipedia.org/wiki/Is_Google_Making_Us_Stupid%3F https://newsletter.pessimistsarchive.org/p/when-educators-mourned-the-slide https://newsletter.pessimistsarchive.org/p/when-educators-mo... New decade, same old argument. It's not > "Claude, think for me" It's > "Claude, be my subordinate and get this done for me" Instead of complaining on the sidelines, I'm getting a shit ton of work done.
- ultrarunner 4mo ago
- metadaemon 4mo agoHas anyone compared a setup like this to just using LM Studio?
- CharlesW 4mo agoYes, I can confirm that LM Studio works great for this.
- hanifbbz 4mo agoHere's a visual post for using LM Studio and VS Code (and Pi): https://blog.alexewerlof.com/p/local-llms-for-agentic-coding https://blog.alexewerlof.com/p/local-llms-for-agentic-coding One way or another local AI is the future. I actually find weaker models more interesting because it keeps me sharp (at the cost of velocity of course).
- deleted 4mo ago[deleted]
- sleepybrett 4mo agoor you can just load up ollama, have it load a local model and point claude or opencode at it... is this article old? It's not. I'm not sure why he went through all the bother of llama.cpp
- malkosta 4mo agoThat was exactly my same question. Then I finished reading the post. The reason is pretty clear, and written in the post: it is faster than ollama+mlx.
- sleepybrett 4mo agohow much faster?
- freerunnering 4mo agoI was benchmarking different models, different engines, and different draft models, I posted a video on twitter, and people started asking about the setup in the final screen recording. So the blog post isn't so much "how a beginner should setup something" it's "here's the setup I posted in the video". Original video: https://x.com/Freerunnering/status/2065275403548168398 https://x.com/Freerunnering/status/2065275403548168398 And in the blog post there is a table showing the different speeds I got from different engines. Slowest combo was 38.1 tk/s, and the fastest was 72.2 tk/s. All from "the same" model.
- malkosta 4mo agoAll I remember is that it's pretty clear written in the post...
- krzyk 4mo agoollama is a wrapper on top of llama.cpp, and it makes llama.cpp slower, why use it? Also Ollama has other issues (like forgetting what it really is - a wrapper).
- flowbarai 4mo ago[flagged]
- rectang 4mo agoDoes anybody run a local agent on a Mac using an outboard GPU?
- benbojangles 4mo agoI run a second Mac for local llm use and access it remotely using ssh from the first mac
- LoganDark 4mo agoI poured a couple days into custom Burn inference for Qwen3-Coder-Next only to find it doesn't come with a speculative decoder, so on my M4 Max I can't push it much further than 120t/s. That's still kinda slow, though still faster than llama.cpp's 70.9t/s and MLX's 80.6t/s with the same model. Claude Fable 5 is recommending I use the Qwen3 MTP -- I worry that will compromise the quality somewhat, but might give it a try to see if I can get more usable speeds.
- reenorap 4mo agoMy biggest pet peeve with all these articles on local AI is the only thing they talk about is tokens per second. No one mentions the quality of the answers. No one. I don't mind waiting a little longer if the quality is better. Quickly serving me slop doesn't make it more useful. Are people really only looking at tokens per second?
- akman 4mo agoThat's fair. There are even many dimensions to define 'quality' which include use case (coding? writing? multimedia?) and prompt. I suppose if you ask testers to provide benchmarks with their analysis, that might hamper their desire to share.
- ozim 4mo agoLocal model as such will give you "autocomplete on steroids" but it is not going to run away and implement cross project feature like frontier model in let's say Cursor. So there is no value in testing quality of answers, but there is value in testing token speed. You just have to have correct expectations.
- krzyk 4mo agoIs autocomplete using LLMs really useful? Even with frontier models I found it to be about 50% right, I turned it of and prefer to use IntelliJ built-in, it is way more reliable. For me local models is all about quality, and how to achieve that - e.g. by providing guardrails that test the job done.
- frollogaston 4mo agoThe model already has its own quality benchmarks elsewhere. The article is just about running the model on X hardware, so the remaining question is then how fast it is. Or does the output quality somehow depend on the hardware too?
- jmkni 4mo agoThe quality is obviously much worse, but still useful as a reference if you generally know what you are doing It solve the "I'm coding on the plane and need to look up this thing I've forgotten" problem, for me at least
- deleted 4mo ago[deleted]
- mark_l_watson 4mo agoNice writeup, thanks. I run something very similar except for directly using pi as the agentic harness I use little-coder that wraps pi with reasonable defaults for running local models. Even though my local setup is a bit slow, it is a thrill to do real work completely locally.
- bicepjai 4mo agoI assumed lmstudio is the obvious choice after ollama. Is there a reason lmstudio is not used widely ?
- stingraycharles 4mo agoYeah I’ve also been using it on macOS, my experience is that it works better with the metal API and has better performance.
- dofm 4mo agoLM Studio is fine. Gorgeous actually. I've found it really helpful for understanding parameters, settings, general figuring out. But there is an incentive not to use it if you want to write an article that uses only open-source tools, because it isn't.
- krzyk 4mo agoWhy would anyone use Ollama at all (aside from obvious reasons one can look up online) - llama.cpp used directly, without this wrapper is faster. Basically one has two real choices for local LLMs: llama.cpp (if single user) or vLLM (if multi-user/enterprise).
- everlier 4mo agoYou can also install Harbor and then it's: harbor up omlx opencode
- anigbrowl 4mo agoThis video is realtime. And shows the agent responding at a perfectly usable speed. Alas, this video appears not have been linked to the text that describes it. Perhaps I should ask an AI to generate an artistic rendering of the author's description.
- freerunnering 4mo agoThe video is stuck in an `<img>` tag so you need to wait for it to load. On a slow connection it might just not show for a while. Though the video is only 1MB so should load in if you wait.
- anigbrowl 4mo agoI have >400MiB/s on this machine and had already spent several minutes reading through the explanation/instructions before scrolling back to the top; it just never loads for me. I had to manually open the link in another tab, for whatever reason.
- jumploops 4mo agoI've been quite impressed with DeepSeek v4 Flash running via antirez's ds4[0]. It feels like a GPT-4 class model in terms of "stored knowledge" but is better at long-horizon tool calling than any of the GPT-4 class models. Running on a 128GB MBP M4 Max, I'm getting ~24 t/s on generation and ~200 t/s on prefill. I was expecting it to feel slow, and it certainly does when e.g. generating code, but it's surprisingly useful as a "machine orchestrator" for simple tasks. For non-agentic usecases, it's a decent enough model to converse with, and has the benefit of being entirely self-contained/private. [0]https://github.com/antirez/ds4 https://github.com/antirez/ds4
- smetannik 4mo agoI wonder why something like LM Studio didn't work for the author?
- b3ing 4mo agoThat’s what I was wondering, lm studio and draw things are easy to use apps that handle much of the cruft for you
- freerunnering 4mo agoI do a lot of fine tuning and development with small models themselves (not just using an LLM over a HTTP API). So downloading the models directly and running them from the CLI was natural for me, so that's what I reached for when I wanted to play around with this.
- teiji-tango 4mo ago[flagged]
- deleted 4mo ago[deleted]
- jlintc 4mo ago[flagged]
- jkwang 4mo ago[flagged]
- datadrivenangel 4mo ago[dead]
- hmontazeri 4mo agoI use LM Studio with the local server it ships and connect it to opencode. Takes 2 min to setup
- godfathermway 4mo ago[dead]
- ljosifov 4mo agoFor high Ram (unified), and relatively middling to lowish Tflops and bandwidth GB/s, usually MoEs are most hopeful. The current top-1 in the (iq, tok/s, @ context depth) ranks for me (M2 Max, 96gb) is DeepSeek-V4-Flash REAP25 <65gb gguf + ds4-server + pi agent. Not better than cloud API ofc, but useful enough to endure if I need to. E.g on a non-Internet 4h flight the battery (local llm draws 60w) held long enough. REAP supporting ds4 branch here https://github.com/ljubomirj/ds4/tree/reap-compact-support https://github.com/ljubomirj/ds4/tree/reap-compact-support DS4F dropping to unusable <10 tok/s only at 784K context (!!) makes a big difference.
- d4rkp4ttern 4mo agoIt’s relatively simple to use llama.cpp/server to spin up a local LLM to work with Claude Code or Codex-CLI. The required llama server settings are often scattered all over so I maintain a set of instructions here for several popular open LLMs: https://pchalasani.github.io/claude-code-tools/integrations/local-llms/ https://pchalasani.github.io/claude-code-tools/integrations/...
- ricardobeat 4mo agoDo you use that as a daily driver? Claude Code' prompt is huge and causes you to spend a long, long time on prompt processing for local models, then running out of context shortly after.
- d4rkp4ttern 4mo agoYes CC prompt can be ~30K tokens. I definitely do not use this as a daily driver. I did use it a few times for sensitive document work with Qwen3.6 MOE.
- deleted 4mo ago[deleted]
- alexwwang 4mo agoI wonder if these local model could really solve problems especially for users that aren’t experts on a given coding language. I am not sure that, more than inline auto completion and unit implementation, are these model capable of designing and composing tech specs that really work.
- new_usemame 4mo ago[flagged]
- bluerooibos 4mo agoI cannot wait until a time in the future when we have local models that are Opus 4.6+ level, and capable of running on inexpensive hardware like a 16Gb Mac. Hopefully that's only a few years away.
- tosief 4mo ago[flagged]
- koliber 4mo agoHow much RAM did the local machine have?
- k2enemy 4mo agoGrammar note: When used as a verb, it should be "set up," and when used as a noun, "setup." Other examples (verb, noun): log in, login back up, backup shut down, shutdown break down, breakdown warm up, warmup
- knightops_dev 4mo ago[flagged]
- yesitcan 4mo agoWhy not just get Claude to set it up?
- zftnb666 4mo ago[flagged]