9 ms·
April 2026 TLDR Setup for Ollama and Gemma 4 26B on a Mac mini
- greenstevester 6mo ago[flagged]
- krzyk 6mo agoBy desk you mean that "Mac mini"? Because it is pricey. In my country it is 1000 USD (from Apple for basic M4 with 24GB). My desk was 1/5th of that price. And considering that this Mac mini won't be doing anything else is there a reason why not just buy subscription from Claude, OpenAI, Google, etc.? Are those open models more performant compared to Sonnet 4.5/4.6? Or have at least bigger context?
- lambda 6mo agoRight now, open models that run on hardware that costs under $5000 can get up to around the performance of Sonnet 3.7. Maybe a bit better on certain tasks if you fine tune them for that specific task or distill some reasoning ability from Opus, but if you look at a broad range of benchmarks, that's about where they land in performance. You can get open models that are competitive with Sonnet 4.6 on benchmarks (though some people say that they focus a bit too heavily on benchmarks, so maybe slightly weaker on real-world tasks than the benchmarks indicate), but you need >500 GiB of VRAM to run even pretty aggressive quantizations (4 bits or less), and to run them at any reasonable speed they need to be on multi-GPU setups rather than the now discontinued Mac Studio 512 GiB. The big advantage is that you have full control, and you're not paying a $200/month subscription and still being throttled on tokens, you are guaranteed that your data is not being used to train models, and you're not financially supporting an industry that many people find questionable. Also, if you want to, you can use "abliterated" versions which strip away the censoring that labs do to cause their models to refuse to answer certain questions, or you can use fine-tunes that adapt it for various other purposes, like improving certain coding abilities, making it better for roleplay, etc.
- zozbot234 6mo agoYou don't need that much VRAM to run the very largest models, these are MoE models where only a small fraction is being computed with at any given time. If you plan to run with multiple GPUs and have enough PCIe lanes (such as with a proper HEDT platform) CPU-GPU transfers start to become a bit less painful. More importantly, streaming weights from disk becomes feasible, which lets you save on expensive RAM. The big labs only avoid this because it costs power at scale compared to keeping weights in DRAM, but that aside it's quite sound.
- lambda 6mo agoWhile you can run with weights in RAM or even disk, it gets a lot slower; even though on any given token a fraction of the weights are used, that can change with each token, so there is a lot of traffic to transfer weights to the GPU, which is a lot slower than if it's directly in GPU RAM. And even more slower if you stream from disk. Possible, yes, and maybe OK for some purposes, but you might find it painfully slow.
- zhongwei2049 6mo agoI have the same setup (M4 Pro, 24GB). The e4b model is surprisingly snappy for quick tasks. The full 26B is usable but not great — loading time alone is enough to break your flow. Re: subscriptions vs local — I use both. Cloud for the heavy stuff, local for when I'm iterating fast and don't want to deal with rate limits or network hiccups.
- redrove 6mo agoThere is virtually no reason to use Ollama over LM Studio or the myriad of other alternatives. Ollama is slower and they started out as a shameless llama.cpp ripoff without giving credit and now they “ported” it to Go which means they’re just vibe code translating llama.cpp, bugs included.
- iLoveOncall 6mo ago> There is virtually no reason to use Ollama over LM Studio or the myriad of other alternatives. Hmm, the fact that Ollama is open-source, can run in Docker, etc.?
- DiabloD3 6mo agoOllama is quasi-open source. In some places in the source code they claim sole ownership of the code, when it is highly derivative of that in llama.cpp (having started its life as a llama.cpp frontend). They keep it the same license, however, MIT. There is no reason to use Ollama as an alternative to llama.cpp, just use the real thing instead.
- simondotau 6mo agoIf it’s MIT code derived from MIT code, in what way is its openness ”quasi”? Issues of attribution and crediting diminish the karma of the derived project, but I don’t see how it diminishes the level of openness.
- DiabloD3 6mo agoFOSS licensing can only exist in terms of Copyright. Without Copyright, you cannot license FOSS. If something has an incorrect Copyright attribution, then the license can be viewed as invalid until this deficiency has been corrected (obv. depending on local laws, etc). On top of this, it would not be unreasonable for the numerous authors of llama.cpp to issue DMCA takedown requests if Ollama is unwilling to correct it.
- alifeinbinary 6mo ago
- easygenes 6mo agoWhy is ollama so many people’s go-to? Genuinely curious, I’ve tried it but it feels overly stripped down / dumbed down vs nearly everything else I’ve used. Lately I’ve been playing with Unsloth Studio and think that’s probably a much better “give it to a beginner” default.
- polotics 6mo agoOllama got some first-mover advantage at the time when actually building and git pulling llama.cpp was a bit of a moat. The devs' docker past probably made them overestimate how much they could lay claim to mindshare. However, no one really could have known how quickly things would evolve... Now I mostly recommend LM-studio to people. What does unsloth-studio bring on top?
- easygenes 6mo agoLM Studio has been around longer. I’ve used it since three years ago. I’d also agree it is generally a better beginner choice then and now. Unsloth Studio is more featureful (well integrated tool calling, web search, and code execution being headline features), and comes from the people consistently making some of the best GGUF quants of all popular models. It also is well documented, easy to setup, and also has good fine-tuning support.
- xenophonf 6mo agoLM Studio isn't free/libre/open source software, which misses the point of using open weights and open source LLMs in the first place.
- vonneumannstan 6mo agoDisagree, there are a lot of reasons to use open source local LLMs that aren't related to free/libre/oss principles. Privacy being a major one.
- 6mo ago
- robotswantdata 6mo agoWhy are you using Ollama? Just use llama.cpp brew install llama.cpp use the inbuilt CLI, Server or Chat interface. + Hook it up to any other app
- Bigsy 6mo agoFor MLX I'd guess.
- redrove 6mo agohttps://omlx.ai/ https://omlx.ai/
- wronglebowski 6mo agoThat also comes upstream from llama.cpp https://github.com/ggml-org/llama.cpp/discussions/4345 https://github.com/ggml-org/llama.cpp/discussions/4345
- boutell 6mo agoLast night I had to install the VO.20 pre-release of ollama to use this model. So I'm wondering if these instructions are accurate.
- logicallee 6mo agoIn case someone would like to know what these are like on this hardware, I tested Gemma 4 32b (the ~20 GB model, the largest Gemma model Google published) and Gemma 4 gemma4:e4b (the ~10 GB model) on this exact setup (Mac Mini M4 with 24 GB of RAM using Ollama), I livestreamed it: https://www.youtube.com/live/G5OVcKO70ns https://www.youtube.com/live/G5OVcKO70ns The ~10 GB model is super speedy, loading in a few seconds and giving responses almost instantly. If you just want to see its performance, it says hello around the 2 minute mark in the video (and fast!) and the ~20 GB model says hello around 5 minutes 45 seconds in the video. You can see the difference in their loading times and speed, which is a substantial difference. I also had each of them complete a difficult coding task, they both got it correct but the 20 GB model was much slower. It's a bit too slow to use on this setup day to day, plus it would take almost all the memory. The 10 GB model could fit comfortably on a Mac Mini 24 GB with plenty of RAM left for everything else, and it seems like you can use it for small-size useful coding tasks.
- aetherspawn 6mo agoWhich harness (IDE) works with this if any? Can I use it for local coding right now?
- lambda 6mo agoYes, you can use it for local coding. Most harnesses can be pointed at a local endpoint which provides an OpenAI compatible API, though I've had some trouble using recent versions of Codex with llama.cpp due to an API incompatibility (Codex uses the newer "responses" API, but in a way that llama.cpp hasn't fully supported). I personally prefer Pi as I like the fact that it's minimalist and extensible. But some people just use Claude Code, some OpenCode, there are a ton of options out there and most of them can be used with local models.
- kristopolous 6mo agoIt needs to support tool calling and many of the quantized ggufs don't so you have to check. I've got a workaround for that called petsitter where it sits as a proxy between the harness and inference engine and emulates additional capabilities through clever prompt engineering and various algorithms. They're abstractly called "tricks" and you can stack them as you please. https://github.com/day50-dev/Petsitter https://github.com/day50-dev/Petsitter You can run the quantized model on ollama, put petsitter in front of it, put the agent harness in front of that and you're good to go If you have trouble, file bugs. Please! Thank you edit: just checked, the ollama version supports everything $ llcat -u http://localhost:11434 -m gemma4:latest --info ["completion", "vision", "audio", "tools", "thinking"] so you can just use that.
- milchek 6mo agoI tested briefly with a MacBook Pro m4 with 36gb. Run in LM Studio with open code as the frontend and it failed over and over on tool calls. Switched back to qwen. Anyone else on similar setup have better luck?
- internet101010 6mo agoI failed to run in LM Studio on M5 with 32gb at even half max context. Literally locked up computer and had to reboot. Ran gemma-4-26B-A4B-it-GGUF:Q4_K_M just fine with llama.cpp though. First time in a long time that I have been impressed by a local model. Both speed (~38t/s) and quality are very nice.
- jasonjmcghee 6mo agoHaven't had time to try yet, but heard from others that they needed to update both the main and runtime versions for things to work.
- abroadwin 6mo agoEven with the latest version of LM Studio and the latest runtimes I find that tool use fails 100% of the time with the following error: Error rendering prompt with jinja template: "Cannot apply filter "upper" to type: UndefinedValue". EDIT: The issue is addressed in LM Studio 0.4.9 (build 1), which auto-update wasn't picking up for me for some reason.
- jasonjmcghee 6mo agoI googled it- supposed fixed template https://github.com/ggml-org/llama.cpp/issues/21347#issuecomment-4182360961 https://github.com/ggml-org/llama.cpp/issues/21347#issuecomm...
- abroadwin 6mo agoAlas, this does not resolve the issue for me.
- 6mo ago
- volume_tech 6mo ago[dead]
- mark_l_watson 6mo agoThe article has a few good tips for using Ollama. Perhaps it should note that the Gemma 4 models are not really trained for strong performance with coding agents like OpenCode, Claude Code, pi, etc. The Gemma 4 models are excellent for applications requiring tool use, data extraction to JSON, etc. I asked Gemini Pro about this earlier and Gemini Pro recommended qwen 3.5 models specifically for coding, and backed that up with interesting material on training. This makes sense, and is something that I do: use strong models to build effective applications using small efficient models.
- Aurornis 6mo ago> I asked Gemini Pro about this earlier and Gemini Pro recommended qwen 3.5 models specifically for coding, and backed that up with interesting material on training. The Gemma models were literally released yesterday. You can’t ask LLMs for advice on these topics and get accurate information. Please don’t repeat LLM-sourced answers as canonical information
- zozbot234 6mo agoIt's not just LLM sourced though, folks have literally tried this after the release with the 26A4B model and it wasn't very good. Maybe the dense ~31B model is worthwhile though.
- Aurornis 6mo agoMany Gemma implementations are or were broken on launch day. The first attempts to fix llama.cpp’s tokenizer were merged hours ago. Everyone hated Qwen3.5 at launch too because so many implementations were broken and couldn’t do tool calling. You need to ignore social media “I tried this and it sucks” echo chambers for new model releases.
- mark_l_watson 6mo agoI agree with your criticism. I should have simply said that I had good results with gemma 4 tool use, and agentic coding with gemma 4 didn’t yet work well for me.
- anonyfox 6mo agoM5 air here with 32gb ram and 10/10 cores. Anyone got some luck with mlx builds on oMLX so far? Not at my machine right now and would love to know if these models already work including tool calling
- smith7018 6mo agoI know that someone got Gemma 4 E4B working with MLX [1] but I don't know much more than that. 1: https://github.com/bolyki01/localllm-gemma4-mlx https://github.com/bolyki01/localllm-gemma4-mlx
- Yukonv 6mo agoThe latest release v0.3.2 has partial support, generation is supported but not all special tokens are handled. I've done some personal testing to add tool calling and <|channel> thinking support. https://github.com/Yukon/omlx https://github.com/Yukon/omlx
- anonyfox 6mo agoawesome man, can’t wait! And just now checked it out and indeed 0.3.2 does already work for baseline chatting with mlx versions of Gemma 4 … downloading and comparing different variants right now!
- renewiltord 6mo agoJust told Claude to sort it out and it ran it. 26 tok/s on the Mac mini I use for personal claw type program. Unusable for local agent but it’s okay.
- zozbot234 6mo agoIsn't 26 tok/s quite usable for a claw-like agent though? You can chat with it on a IM platform and get notified as soon as it replies, you're not dependent on real-time quick interaction.
- renewiltord 6mo agoFor me it's too slow. Prefer using cloud agent. Can do more tasks.
- Aurornis 6mo agoIf this is your first time using open weight models right after release, know that there are always bugs in the early implementations and even quantizations. Every project races to have support on launch day so they don’t lose users, but the output you get may not be correct. There are already several problems being discovered in tokenizer implementations and quantizations may have problems too if they use imatrix. So you’re going to see a lot of “I tried it but it sucks because it can’t even do tool calls” and other reports about how the models don’t work at all in the coming weeks from people who don’t realize they were using broken implementations. If you want to try cutting edge open models you need to be ready to constantly update your inference engine and check your quantization for updates and re-download when it’s changed. The mad rush to support it on launch day means everything gets shipped as soon as it looks like it can produce output tokens, not when it’s tested to be correct.
- colechristensen 6mo agoYou seem like you know what you're talking about... what inference engine should I use? (linux, 4090) I keep having "I tried it but it sucks" issues mostly around tool calling and it's not clear if it's the model or ollama. And not one model in particular, any of them really.
- vardalab 6mo agojust use openrouter or google ai playground for the first week till bugs are ironed out. You still learn the nuances of the model and then yuu can switch to local. In addition you might pickup enough nuance to see if quantization is having any effect
- Aurornis 6mo agoI don’t know if any of engines are fully tested yet. For new LLMs I get in the habit of building llama.cpp from upstream head and checking for updated quantizations right before I start using it. You can also download llama.cpp CI builds from their release page but on Linux it’s easy to set up a local build. If you don’t want to be a guinea pig for untested work then the safe option would be to wait 2-3 weeks
- kristopolous 6mo agoAre you getting tool call and multimodal working? I don't see it in the quantized unsloth ggufs...
- zachperkel 6mo agohow many TPS does a build like this achieve on gemma 4 26b?
- kanehorikawa 6mo ago[dead]
- neo_doom 6mo agoHuge Claude user here… can someone help me set some realistic expectations if I bought a Mac mini and spun one up? I use Claude primarily for dev work and Home Lab projects. Are the open models good enough to run locally and replace the Claude workload? Or am I better off with my $20/mo Claude subscription?
- NietTim 6mo agoThey are good for small tasks but you would not be able to use it like you use Claude and most likely be disappointed. But also, I do not know how you use claude. There are many services online which offer hosted services for these models, my advice for anyone who is thinking about buying hardware to self host this is to try those first, that way you can get an impression of the capabilities and limitations of those models before you commit to buying hardware
- alfiedotwtf 6mo agoSo far, I’ve found gpt-oss-20B to be pretty good agentic wise, but it’s nothing like Claude Code using its paid models. (I haven’t tried the 120B, which I’ve read is significantly better than 20B)
- hamdingers 6mo agoBest way to find out is to buy $10 of OpenRouter credits and try the models for yourself. From my experience doing this, they're nowhere close, but it's entertaining to check in once in a while.
- MrScruff 6mo agoI've been playing with the open models since the original llama leak. They're getting better over time, are useful for tasks of moderate complexity and it's just cool to have a binary blob of knowledge that you can run locally without an internet connection. However you should manage your expectations. Whatever the benchmarks say, you'll quickly realise they're not at all competing with Sonnet let alone Opus. Even the largest open weights models aren't really doing that.
- techpulselab 6mo ago[dead]
- jiusanzhou 6mo ago[dead]
- spencer-p 6mo agoWeird that the steps are for "Gemma 4 12b", which does not exist, and then switches to 26b midway through. There's also a step to verify that it doesn't fit on the GPU with ollama ps showing "14%/86% CPU/GPU". Doesn't this mean you'll have really bad performance?
- Schiendelman 6mo agoThe Mac mini doesn't have different memory for the CPU and GPU, so maybe that's ignorable?
- aplomb1026 6mo ago[dead]
- jasonriddle 6mo agoSlightly off topic, but question for folks. I'm hoping to replace coding with Claude Sonnet 4.5 with a model with an open source or open weights model. Are any of the models on Ollama.com cloud offering (https://ollama.com/search?c=cloud https://ollama.com/search?c=cloud) or any of the models on OpenRouter.ai a close replacement? I know that no model right now matches the full performance and capabilities of Claude Sonnet 4.5, but I want to know how close I can get and with which model(s). If there is a model you say can replace it, talk about how long you have been using it for, and using what harness (Claude code, opencode, etc), and some strengths and weakness you have noticed. I'm not interested in what benchmarks say, I want to hear about real world use from programmers using these models.
- scottcha 6mo agoYes GLM5 and KimiK2.5 are pretty close replacements for sonnet.
- jasonriddle 6mo agoWhat coding harness are you using? What are some example workflows you have used either for? Have you used them only for new/simple projects or for more complicated refactoring or architecture design?
- scottcha 6mo agoI use OpenCode and have just started using Nanoclaw with ClaudeCode (my coworker has a post coming on this) and sometimes ClaudeCode with Claude Code Router. I do a range of small to complex work with these but I also do drop back in to Claude Opus for some really complex things where I want it to be more autonomous.
- MrScruff 6mo agoHaven't really tried GLM5 much but I've used 4.7 quite a bit and it was pretty far from competing with Sonnet at the time, although I saw claims online to the contrary.
- 6mo ago
- OkGoDoIt 6mo agoSorry for being off topic, but why can’t I open this without being logged into GitHub? I thought gists are either completely private or publicly accessible. Are they no longer publicly accessible?
- OkGoDoIt 6mo agoIn case anyone’s wondering, I tried it again and it worked this time, even without logging in. Maybe because this was my first visit to GitHub in a new country (I’m currently on vacation), I triggered some sort of anti-scraping measure or something.
- kilzimir 6mo agoKinda crazy that I can run a 26B model on a 1500€ laptop (MacBook Air M5 32GB). Does anyone know how I can actually use this in a productive way?
- pwr1 6mo agoRunning 26B locally is impressive but the latency math gets rough once your doing anything beyond chat. We switched from local inference to API calls for image generation specifically because cold start + generation time on consumer hardware made it impractical for any kind of automated workflow. Local is great for experimentation but production workloads that need to run reliably at specific times still favor API imo. That said for privacy sensitive use cases where data cant leave the machine, setups like this are invaluable.
- aplomb1026 6mo ago[dead]
- amelius 6mo agoHas anyone tried to run it on a Jetson Orin AGX with 64GB unified memory?
- Xentyon 6mo agoNice setup. Running models locally on Mac hardware has gotten surprisingly viable. I'm using a similar stack in Switzerland for testing AI agent workflows — the M-series chips handle inference well for tool-calling tasks.
- aimemobe 6mo ago[flagged]