33 ms·
Ollama is now powered by MLX on Apple Silicon in preview
- jedisct1 6mo agoWorks really great with https://swival.dev https://swival.dev and qwen3.5.
- babblingfish 6mo agoLLMs on device is the future. It's more secure and solves the problem of too much demand for inference compared to data center supply, it also would use less electricity. It's just a matter of getting the performance good enough. Most users don't need frontier model performance.
- gedy 6mo agoMan I really hope so, as, as much as I like Claude Code, I hate the company paying for it and tracking your usage, bullshit management control, etc. I feel like I'm training my replacement. Things feel like they are tightening vs more power and freedom. On device I would gladly pay for good hardware - it's my machine and I'm using as I see fit like an IDE.
- susupro1 6mo ago[dead]
- aurareturn 6mo agoWhen local LLMs get good enough for you to use delightfully, cloud LLMs will have gotten so much smarter that you'll still use it for stuff that needs more intelligence.
- gedy 6mo agoTrue, but I'm already producing code/features faster than company knows what to do with, (even though every company says "omg we need this yesterday", etc). Even coding before AI was basically same. Code tools that free my time up is very nice.
- dgb23 6mo agoThat's not necessarily the case. So far, commercial cloud LLMs have maintained a head-start, but there is no law of nature that prevents us from having competitive open models. In fact the space seems to move at a rapid pace as more and more specialized models come out. There's a possible trajectory where open weight models will compete side by side or even be preferable for many use cases, just like what happened with OS's and SQL DB's.
- aurareturn 6mo agoIt isn't going to replace cloud LLMs since cloud LLMs will always be faster in throughput and smarter. Cloud and local LLMs will grow together, not replace each other. I'm not convinced that local LLMs use less electricity either. Per token at the same level of intelligence, cloud LLMs should run circles around local LLMs in efficiency. If it doesn't, what are we paying hundreds of billions of dollars for? I think local LLMs will continue to grow and there will be an "ChatGPT" moment for it when good enough models meet good enough hardware. We're not there yet though. Note, this is why I'm big on investing in chip manufacture companies. Not only are they completely maxed out due to cloud LLMs, but soon, they will be double maxed out having to replace local computer chips with ones that are suited for inferencing AI. This is a massive transition and will fuel another chip manufacturing boom.
- AugSun 6mo agoLooking at downvotes I feel good about SDE future in 3-5 years. We will have a swamp of "vibe-experts" who won't be able to pay 100K a month to CC. Meanwhile, people who still remember how to code in Vim will (slowly) get back to pre-COVID TC levels.
- QuantumNomad_ 6mo agoWhat is CC and TC? I have not heard these abbreviations (except for CC to mean credit card or carbon copy, neither of which is what I think you mean here).
- Ericson2314 6mo agoI figured it out from context clues CC: Claude Code TC: total comp(ensation)
- AugSun 6mo agoThank you for clarifying! (I had no idea it needs to be explained, sorry.)
- deleted 6mo ago
- AugSun 6mo ago"Most users don't need frontier model performance" unfortunately, this is not the case.
- AugSun 6mo ago[flagged]
- seanhunter 6mo agoComplaining about downvotes is futile and is also against hn guidelines.
- AugSun 6mo agoI'm not complaining "about downvotes" LOL I'm explaining why some people will be replaced by LLMs because of their own "context window" length.
- selcuka 6mo agoAny citations? Because that was my impression, too. I want frontier model performance for my coding assistant, but "most users" could do with smaller/faster models. ChatGPT free falls back to GPT-5.2 Mini after a few interactions.
- asutekku 6mo agoFrontier model has much better knowledge and they usually hallucinate less. It's not about the coding capabilities, it's about how much you can trust the model.
- Barbing 6mo agore: trust- Have you tried the free version of ChatGPT? It is positively appalling. It’s like GPT 3.5 but prompted to write three times as much as necessary to seem useful. I wonder how many people have embarrassed themselves, lost their jobs, and been critically misinformed. All easy with state-of-the-art models but seemingly a guarantee with the bottom sub-slop tier. Is the average person just talking to it about their day or something?
- melvinroest 6mo agoI have journaled digitally for the last 5 years with this expectation. Recently I built a graphRAG app with Qwen 3.5 4b for small tasks like classifying what type of question I am asking or the entity extraction process itself, as graphRAG depends on extracted triplets (entity1, relationship_to, entity2). I used Qwen 3.5 27b for actually answering my questions. It works pretty well. I have to be a bit patient but that’s it. So in that particular use case, I would agree. I used MLX and my M1 64GB device. I found that MLX definitely works faster when it comes to extracting entities and triplets in batches.
- nkzd 6mo agoDid you get any insights about yourself from this process? I am thinking of doing the same
- melvinroest 6mo agoTL;DR: you don't need to do any treasure hunt on your notes by just typing stuff into the search bar. Having your own graphRAG system + LLM on your notes is basically a "Google" but then on your own notes. Any question you have: if you have a note for it, it will bubble up. The annoying thing is that false positives will also bubble up. ---- Full reaction: Yes but perhaps not in a way you might expect. Qwen's reasoning ability isn't exactly groundbreaking. But it's good enough to weave a story, provided it has some solid facts or notes. GraphRAG is definitely a good way to get some good facts, provided your notes are valuable to you and/or contain some good facts. So the added value is that you now have a super charged information retrieval system on your notes with an LLM that can stitch loose facts reasonably well together, like a librarian would. It's also very easy to see hallucinations, if you recognize your own writing well, which I do. The second thing is that I have a hard time rereading all my notes. I write a lot of notes, and don't have the time to reread any of them. So oftentimes I forget my own advice. Now that I have a super charged information retrieval system on my notes, whenever I ask a question: the graphRAG + LLM search for the most relevant notes related to my question. I've found that 20% of what I wrote is incredibly useful and is stuff that I forgot. And there are nuggets of wisdom in there that are quite nuanced. For me specifically, I've seen insights in how I relate to work that I should do more with. I'll probably forget most things again but I can reuse my system and at some point I'll remember what I actually need to remember. For example, one thing I read was that work doesn't feel like work for me if I get to dive in, zoom out, dive in, zoom out. Because in the way I work as a person: that means I'm always resting and always have energy for the task that I'm doing. Another thing that it got me to do was to reboot a small meditation practice by using implementation intentions (e.g. "if I wake up then I meditate for at least a brief amount of time"). What also helps is to have a bit of a back and forth with your notes and then copy/paste the whole conversation in Claude to see if Claude has anything in its training data that might give some extra insight. It could also be that it just helps with firing off 10 search queries and finds a blog post that is useful to the conversation that you've had with your local LLM.
- pezgrande 6mo agoYou could argue that the only reason we have good open-weight models is because companies are trying to undermine the big dogs, and they are spending millions to make sure they dont get too far ahead. If the bubble pops then there wont be incentive to keep doing it.
- aurareturn 6mo agoI agree. I can totally see in the future that open source LLMs will turn into paying a lumpsum for the model. Many will shut down. Some will turn into closed source labs. When VCs inevitably ask their AI labs to start making money or shut down, those free open source LLMS will cease to be free. Chinese AI labs have to release free open source models because they distill from OpenAI and Anthropic. They will always be behind. Therefore, they can't charge the same prices as OpenAI and Anthropic. Free open source is how they can get attention and how they can stay fairly close to OpenAI and Anthropic. They have to distill because they're banned from Nvidia chips and TSMC. Before people tell me Chinese AI labs do use Nvidia chips, there is a huge difference between using older gimped Nvidia H100 (called H20) chips or sneaking around Southeast Asia for Blackwell chips and officially being allowed to buy millions of Nvidia's latest chips to build massive gigawatt data centers.
- spiderfarmer 6mo ago“They will always be behind” Car manufacturers said the same.
- aurareturn 6mo agoIt did take decades to catch and surpass US car makers right?
- seanmcdirmid 6mo agoAbout 2.5 decades from the start of the JVs, but they did it. Semiconductors and jet turbines are really the last two tech trees that China has yet to master.
- deleted 6mo ago[deleted]
- karimf 6mo agoDepending on the use case, the future is already here. For example, last week I built a real-time voice AI running locally on iPhone 15. One use case is for people learning speaking english. The STT is quite good and the small LLM is enough for basic conversation. https://github.com/fikrikarim/volocal https://github.com/fikrikarim/volocal
- Barbing 6mo agoBrilliant. Hope to see you in the App Store!
- podlp 6mo agoThat’s awesome! I’ve got a similar project for macOS/ iOS using the Apple Intelligence models and on-device STT Transcriber APIs. Do you think it the models you’re using could be quantized more that they could be downloaded on first run using Background Assets? Maybe we’re not there yet, but I’m interested in a better, local Siri like this with some sort of “agentic lite” capabilities.
- karimf 6mo ago> Do you think it the models you’re using could be quantized more that they could be downloaded on first run using Background Assets? I first tried the Qwen 3.5 0.8B Q4_K_S and the model couldn't hold a basic conversation. Although I haven't tried lower quants on 2B. I'm also interested on the Apple Foundation models, and it's something I plan to try next. AFAIK it's on par with Qwen-3-4B [0]. The biggest upside as you alluded to is that you don't need to download it, which is huge for user onboarding. [0] https://machinelearning.apple.com/research/apple-foundation-models-2025-updates https://machinelearning.apple.com/research/apple-foundation-...
- troad 6mo agoI very recently installed llama.cpp on my consumer-grade M4 MBP, and I've been having loads of fun poking and prodding the local models. There's now a ChatGPT style interface baked into llama.cpp, which is very handy for quick experimentation. (I'm not entirely sure what Ollama would get me that llama.cpp doesn't, happy to hear suggestions!) There are some surprisingly decent models that happily fit even into a mere 16 gigs of RAM. The recent Qwen 3.5 9B model is pretty good, though it did trip all over itself to avoid telling me what happened on Tiananmen Square in 1989. (But then I tried something called "Qwen3.5-9B-Uncensored-HauhauCS-Aggressive", which veers so hard the other way that it will happily write up a detailed plan for your upcoming invasion of Belgium, so I guess it all balances out?)
- whackernews 6mo agoOh does llama.cpp use MLX or whatever? I had this question, wonder if you know? A search suggests it doesn’t but I don’t really understand.
- irusensei 6mo ago>Oh does llama.cpp use MLX or whatever? No. It runs on MacOS but uses Metal instead of MLX.
- zozbot234 6mo agoANE-powered inference (at least for prefill, which is a key bottleneck on pre-M5 platforms) is also in the works, per https://github.com/ggml-org/llama.cpp/issues/10453#issuecomment-4148905254 https://github.com/ggml-org/llama.cpp/issues/10453#issuecomm...
- OkGoDoIt 6mo agoIs that better or worse?
- irusensei 6mo agoDepends. MLX is faster because it has better integration with Apple hardware. On the other hand GGUF is a far more popular format so there will be more programs and model variety. So its kinda like having a very specific diet that you swear is better for you but you can only order food from a few restaurants.
- overfeed 6mo ago> It's just a matter of getting the performance good enough. Who will pay for the ongoing development of (near-)SoTA local models? The good open-weight models are all developed by for-profit companies - you know how that story will end.
- DrScientist 6mo agoApple via customers paying for the whole solution ( eg a laptop that can run decent local models )? I think Apple had something in the region of 143 billion in revenue in the last quarter. Not saying it will happen - just that there are a variety of business models out there and in the end it all depends on where consumers put their money.
- nikanj 6mo agoThat also means sending every user a copy of the model that you spend billions training. The current model (running the models at the vendor side) makes it much easier to protect that investment
- jl6 6mo agoNot sure about the using less electricity part. With batching, it’s more efficient to serve multiple users simultaneously.
- TeMPOraL 6mo agoIndeed. Data centers have so many ways and reasons to be much more energy-efficient than local compute it's not even funny.
- chongli 6mo agoThey do, though I don’t think they max out on energy efficient technology. It’s much easier to cut a deal for cheap electricity with a regional government, much to the chagrin of the locals (who see their power bills go up).
- ZeroGravitas 6mo agoIt feels like you'll soon need a local llm to intermediate with the remote llm, like an ad blocker for browsers to stop them injecting ads or remind you not to send corporate IP out onto the Internet.
- tomashubelbauer 6mo agoI'd like to coin the term "user agent" for this
- blitzar 6mo ago"copilot" seems a good term could also be considered a triage layer
- miki123211 6mo ago> would use less electricity Sorry to shatter your bubble, but this is patently false, LLMs are far more efficient on hardware that simultaneously serves many requests at once. There's also the (environmental and monetary) cost of producing overpowered devices that sit idle when you're not using them, in contrast to a cloud GPU, which can be rented out to whoever needs it at a given moment, potentially at a lower cost during periods of lower demand. Many LLM workloads aren't even that latency sensitive, so it's far easier to move them closer to renewable energy than to move that energy closer to you.
- kortilla 6mo agoWell this is an article about running on hardware I already have in my house. In the winter that’s just a little extra electricity that converts into “free” resistive heating.
- ysleepy 6mo agoI'm actually not sure that's true. Apart from people buying the device with or without the neural accelerator, the perf/watt could be on par or better with the big iron. The efficiency sweet-spot is usually below the peak performance point, see big.little architectures etc.
- zozbot234 6mo ago> LLMs are far more efficient on hardware that simultaneously serves many requests at once. The LLM inference itself may be more efficient (though this may be impacted by different throughput vs. latency tradeoffs; local inference makes it easier to run with higher latency) but making the hardware is not. The cost for datacenter-class hardware is orders of magnitude higher, and repurposing existing hardware is a real gain in efficiency.
- Tepix 6mo agoSeems doubtful. The utilisation will be super high for data center silicon whereas your PC or phone at home is mostly idle.
- thih9 6mo ago> it also would use less electricity How would it use less electricity? I’d like to learn more.
- jychang 6mo agoThat's completely not true. LLM on device would use MORE electricity. Service providers that do batch>1 inference are a lot more efficient per watt. Local inference can only do batch=1 inference, which is very inefficient.
- amelius 6mo agoLLM in silicon is the future. It won't be long until you can just plug an LLM chip into your computer and talk to it at 100x the speed of current LLMs. Capability will be lower but their speed will make up for it.
- theshrike79 6mo agoI'm expecting someone to come up with an LLM version of the Coral USB Accelerator: https://www.coral.ai/products/accelerator https://www.coral.ai/products/accelerator Just plug in a stick in your USB-C port or add an M.2 or PCIe board and you'll get dramatically faster AI inference.
- angoragoats 6mo agoI think there are drastic differences between computer vision models and LLMs that you’re not considering. LLMs are huge relative to vision models, and require gobs of fast memory. For this reason a little USB dongle isn’t going to cut it. Put another way, there already exist add-in boards like this, and they’re called GPUs.
- amelius 6mo agoGPUs are still software programmable. An "LLM chip" does not need that and so can be much more efficient.
- angoragoats 6mo agoSure, but that’s somewhat orthogonal to the point I was making, which is that LLMs are huge in size. Even in the case of a custom “LLM chip,” you’ll need huge amounts of very fast storage of some sort (likely DRAM), which places constraints on the size, power consumption, and cost of such a device. This device, if it existed, would not in any way resemble the Coral TPU product that the GP was referencing; I think in fact it would be closer in size, price, and form factor to a GPU.
- jillesvangurp 6mo ago
- zozbot234 6mo ago> Most users don't need frontier model performance. SSD weights offload makes it feasible to run SOTA local models on consumer or prosumer/enthusiast-class platforms, though with very low throughput (the SSD offload bandwidth is a huge bottleneck, mitigated by having a lot of RAM for caching). But if you only need SOTA performance rarely and can wait for the answer, it becomes a great option.
- iNic 6mo agoIt will probably be a future. My guess is that for many businesses it will still make sense to have more powerful models and to run them centralized in a datacenter. Also, by batching queries you can get efficiencies at scale that might be hard to replicate locally. I can also see a hybrid approach where local models get good at handing off to cloud models for complex queries.
- niek_pas 6mo ago> For many businesses it will still make sense to have more powerful models and to run them centralized in a datacenter. Agree, and I think of it this way: for a lot of businesses, it already makes sense to have a bunch of more powerful computers and run them centralized in a datacenter. Nevertheless, most people at most companies do most of their work on their Macbook Air or Dell whatever. I think LLMs will follow a similar pattern: local for 90% of use cases, powerful models (either on-site in a datacenter or via a service) for everything else.
- goldenarm 6mo agoIt's more secure, but it would make supply much much worse. Data centers use GPU batching, much higher utilisation rates, and more efficient hardware. It's borderline two order of magnitude more efficient than your desktop.
- nbenitezl 6mo agoBut when using it on the cloud a LLM can consult 50 websites, which is super fast for their datacenters as they are backbone of internet, instead you'll have to wait much more on your device to consult those websites before giving you the LLM response. Am i wrong?
- comboy 6mo agoAs things stand today even when doing research tasks, time spent by model is >> than fetching websites. I don't see it changing any time soon, except when some deals happen behind the scenes where agents get to access CF guarded resources that normally get blocked from automated access.
- Const-me 6mo agoWhile data centres indeed have awesome internet connectivity, don’t forget the bandwidth is shared by all clients using a particular server. If you have 100 mbit/sec internet connection at home, a computer in a data centre has 10 gbit/sec, but the server is serving 200 concurrent clients — your bandwidth is twice as fast.
- dwayne_dibley 6mo agoThis might be how Apple will start to see even more sales, the M series processors are so far ahead of anything else, local LLMs could be their main selling point.
- 3yr-i-frew-up 6mo ago[dead]
- konschubert 6mo agoI disagree with every sentence of this. > solves the problem of too much demand for inference False, it creates consumer demand for inference chips, which will be badly utilised. > also would use less electricity What makes you think that? (MAYBE you can save power on cooling. But not if the data center is close to a natural heat sink) > It's just a matter of getting the performance good enough. The performance limitations are inherent to the limited compute and memory. > Most users don't need frontier model performance. What makes you think that?
- deleted 6mo ago[deleted]
- ekianjo 6mo ago> What makes you think that? Looking at actual users of LLMs
- konschubert 6mo agoWhile not everybody is a professional in YOUR domain, many people are professionals in SOME domain. And even outside of that, they deserve a smart conversation partner, for example on topics like health and politics.
- locknitpicker 6mo ago> What makes you think that? The fact that today's and yesterday's models are quite capable of handling mundane tasks, and even companies behind frontier models are investing heavily in strategies to manage context instead of blindly plowing through problems with brute-force generalist models. But let's flip this around: what on earth even suggests to you that most users need frontier models?
- konschubert 6mo agoEverybody has difficult decisions to make in their daily lives and in their work. Having access to a model that is drawing from good sources and takes time to think instead of hallucinating a response is important in many domains of life.
- g947o 6mo agoHave you spent more than 10 min actually running LLM on a local machine? As it stands today, local LLMs don't work remotely as well as some people try to picture them, in almost every way -- speed, performance, cost, usability etc. The only upside is privacy.
- RALaBarge 6mo agoI agree with you in the sense that if you tried to take any model right now and cram it into an iphone, it wouldnt be a claude-level agent. I run 32b agents locally on a big video card, and smaller ones in CPU, but the lack there isn't the logic or reasoning, it is the chain of tooling that Claude Code and other stacks have built in. Doing a lot of testing recently with my own harness, you would not believe the quality improvement you can get from a smaller LLM with really good opening context. Even Microsoft is working on 1-bit LLMs...it sucks right now, but what about in 5 years? But the OP is correct -- everything will have an LLM on it eventually, much sooner than people who do not understand what is going on right now would ever believe is possible.
- kylehotchkiss 6mo agoYes. I've spent months running Qwen2.5-8B on my barebones 16gb ram M4 Mac mini to handle identifying sites from google search results. It has been rock solid. I'm not even running this MLX-powered improvement on it yet. Your idea of what people need from Local LLMs and others are different. Not everybody needs a /r/myboyfriendisai level performance.
- g947o 6mo agoYou probably want to double check the comment I was responding to.
- adam_patarino 6mo ago[flagged]
- podlp 6mo agoRig sounds cool, I just joined the waitlist! I’m building something similar although with a much narrower purpose. Excited to learn more
- adam_patarino 6mo agoTell me more! Thanks for the waitlist
- podlp 6mo agoSent a LinkedIn request. I’m building a language-specific coding agent using Apple Intelligence with custom adapters. It’s more a proof-of-concept at this point, but basic functionality actually works! The 4K context window is brutal, but there’s a variety of techniques to work around it. Tighter feedback loops, linters, LSPs, and other tools to vet generated code. Plus mechanisms for on-device or web-based API discovery. My hypothesis is if all this can work “well enough” for one language/ runtime, it could be adapted for N languages/ runtimes.
- eeixlk 6mo agoObviously apple would prefer this. It would boost demand for more powerful and expensive devices, and align with their privacy marketing. But they have massively fumbled with siri for a long time and then missed huge deadlines with ai promises. Despite having billions, they have shown no competency in delivering services or accurately marketing what to expect from ai features.
- jonhohle 6mo agoI’ve been using google search AI and Gemini, which I find generally pretty good. In the past week, Gemini and Search AI have been bringing in various details of previous searches I’ve done and Search AI conversations I’ve had and it’s extremely gross and creepy. I was looking for details about cars and it started interjecting how the safety would affect my children by name in a conversation where I never mention my children. I was asking details about Thunderbolt and modern Ryzen processors and a fresh Gemini chat brought in details about a completely unrelated project I work on. I’ve always thought local LLMs would be important, but whatever Google did in the past few weeks has made that even more clear.
- theChaparral 6mo agoIt's Personal Intelligence in the Gemini settings. I just turned that off last night when it was doing similar things.
- Aurornis 6mo ago> solves the problem of too much demand for inference compared to data center supply Maybe in the distant future when device compute capacity has increased by multiples and efficiency improvements have made smaller LLMs better. The current data center buildouts are using GPU clusters and hybrid compute servers that are so much more powerful than anything you can run at home that they’re not in the same league. Even among the open models that you can run at home if you’re willing to spend $40K on hardware, the prefill and token generation speeds are so slow compared to SOTA served models that you really have to be dedicated to avoiding the cloud to run these. We won’t be in a data center crunch forever. I would not be surprised if we have a period of data center oversupply after this rush to build out capacity. However at the current rate of progress I don’t see local compute catching up to hosted models in quality and usability (speed) before data center capacity catches up to demand. This is coming from someone who spends more than is reasonable on local compute hardware.
- babblingfish 6mo agoI see a lot of people are confused about the electricity claim so I'll elaborate on it more. The assumption I'm making here is that on device people will run smaller models, that can fit on their machines without needing to buy new computers. If everyone ran inference on their machine there would be no need for these massive datacenters which use huge quantities of electricity. It would utilize the machines they already have and the electricity they're already using. People are making a comparison of the cost per inference or token or whatever and saying datacenters are more efficient which makes obvious sense. What i'm saying is if we eliminate the need for building out dozens of gigawatt datacenters completely then we would use less electricity. I feel like this makes intuitive sense. People are getting lost in the details about cost per inference, and performance on different models.
- codelion 6mo agoHow does it compare to some of the newer mlx inference engines like optiq that support turboquantization - https://mlx-optiq.pages.dev/ https://mlx-optiq.pages.dev/
- dial9-1 6mo agostill waiting for the day I can comfortably run Claude Code with local llm's on MacOS with only 16gb of ram
- gedy 6mo agoHow close is this? It says it needs 32GB min?
- HDBaseT 6mo agoYou can run Qwen3.5-35B-A3B on 32GB of RAM sure, although to get 'Claude Code' performance, which I assume he means Sonnet or Opus level models in 2026, this will likely be a few years away before its runnable locally (with reasonable hardware).
- Foobar8568 6mo agoI fully agree, I run that one with Q4 on my MBP, and the performance (including quality of response) is a let down. I am wondering how people rave so much about local "small devices" LLM vs what codex or Claude code are capable of. Sadly there are too much hype on local LLM, they look great for 5min tests and that's it.
- brcmthrowaway 6mo agoJust train it better with AGENTS.md
- Hamuko 6mo agoI'm reading "more than 32GB of unified memory" to mean at least a 36 GB model.
- rubymamis 6mo agoDoesn't OpenCode supports local models?
- 6mo ago
- LuxBennu 6mo agoAlready running qwen 70b 4-bit on m2 max 96gb through llama.cpp and it's pretty solid for day to day stuff. The mlx switch is interesting because ollama was basically shelling out to llama.cpp on mac before, so native mlx should mean better memory handling on apple silicon. Curious to see how it compares on the bigger models vs the gguf path
- zozbot234 6mo agoThey initially messed up this launch and overwrote some of the GGUF models in their library, making them non-downloadable on platforms other than Apple Silicon. Hopefully that gets fixed.
- goldenarm 6mo agoHow many tokens per second?
- LuxBennu 6mo agoRoughly 8-12 token/s on generation depending on context length. Prompt processing is faster obviously. Haven't benchmarked it super carefully though, just eyeballing the llama.cpp output.
- yg1112 6mo agoThe key difference is that MLX's array model assumes unified memory from the ground up. llama.cpp's Metal backend works fine but carries abstractions from the discrete GPU world — explicit buffer synchronization, command buffer boundaries — that are unnecessary when CPU and GPU share the same address space. You'll notice the gap most at large context lengths where KV cache pressure is highest.
- lioeters 6mo agoInsightful comment, thanks!
- LuxBennu 6mo agothat tracks with what i've noticed practically. shorter prompts feel basically the same between llama.cpp metal and what i'd expect from native mlx, but once context gets longer the overhead starts showing up. would be interesting to see if ollama's mlx path actually handles kv cache differently under the hood or if it just skips the buffer sync layer
- AugSun 6mo ago"We can run your dumbed down models faster": #The use of NVFP4 results in a 3.5x reduction in model memory footprint relative to FP16 and a 1.8x reduction compared to FP8, while maintaining model accuracy with less than 1% degradation on key language modeling tasks for some models.
- brcmthrowaway 6mo agoWhat is the difference between Ollama, llama.cpp, ggml and gguf?
- xiconfjs 6mo agoOllama on MacOS is a one-click solution with stable obe-click updates. Happy so far. But the mlx support was the only missing piece for me.
- yard2010 6mo agoCan you please write about your hardware?
- xiconfjs 6mo ago* macOS 26.x on MacBookPro M1 Max 32GB * Ollama on macOS, cursor to play around * Open WebUI [1] on my Homeserver via API to Ollama (also for remote „A.I.“ access) * running gpt-oss:20b, qwen3.5:9b with ease, qwen3.5:27b for more complex tasks [1] https://github.com/open-webui/open-webui https://github.com/open-webui/open-webui
- brcmthrowaway 6mo agoSeems complicated. Switch to LMStudio
- xiconfjs 6mo agoI tried man times but at least with its API active, LMStudio has some kind of memory leaks which will slow down the whole system (after ~1-2 days of uptime) even after unloading the model and stopping LMStudio up to a point where even playing a 1080p video results in frame drops. No such issues with Ollama.
- benob 6mo agoOllama is a user-friendly UI for LLM inference. It is powered by llama.cpp (or a fork of it) which is more power-user oriented and requires command-line wrangling. GGML is the math library behind llama.cpp and GGUF is the associated file format used for storing LLM weights.
- mfa1999 6mo agoHow does this compare to llama.cpp in terms of performance?
- solarkraft 6mo agoMLX is a bit faster (low double digit percentage), but uses a bit more RAM. Worthwhile tradeoff for many.
- ysleepy 6mo agoOn my M4 Pro MLX has almost 2x tok/s
- firekey_browser 6mo ago[dead]
- deleted 6mo ago[deleted]
- puskuruk 6mo agoFinally! My local infra is waiting for it for months!
- Yukonv 6mo agoGood to see Ollama is catching up with the times for inference on Mac. MLX powered inference makes a big difference, especially on M5 as their graphs point out. What really has been a game changer for my workflow is using https://omlx.ai/ https://omlx.ai/ that has SSD KV cold caching. No longer have to worry about a session falling out of memory and needing to prefill again. Combine that with the M5 Max prefill speed means more time is spend on generation than waiting for 50k+ content window to process.
- davesque 6mo agoYeah omlx seems to me like the front runner right now for running MLX models locally in agent workflows (which depend heavily on caching).
- charlotte12345 6mo ago[dead]
- charlotte12345 6mo ago[flagged]
- universa1 6mo agoi am curious: is the performance gap between x86 cpu inference and apple silicon, or, a imho more apples-to-apples comparison, e.g., amd strixpoint halo vs apple silicon? i would expect the "pure" cpu inference to be behind, but an approach like strix halo/dgx spark to be much closer?
- robotswantdata 6mo agoWhy are people still using Ollama? Serious. Lemonade or even llama.cpp are much better optimised and arguably just as easy to use.
- vorticalbox 6mo agoi like ollama, mostly because the cli is pretty nice. its desktop app has stupid choices like if a model can support tools then the ui should give me the "search" option but it only shows for cloud models. i have ran lmstudio for a while but i don't really use local models that much other than to mess about.
- zozbot234 6mo agoYou can also use OpenWebUI locally which should give you a nice friendly UX once you set it up.
- niek_pas 6mo agoSerious answer: I don't use it that much, it's what I happened to download like 1.5 years ago, and it works fine. Happy to see what may be a speed boost, and have little interest in switching to something else (unless my situation changes, of course).
- eddieroger 6mo ago`ollama serve` and `ollama run` The devex is great and familiar to folks who have used Docker. Reading through the Lemonade documentation, it seems like a natural migration, but we're talking about two steps for getting started versus just one. So I'd need a reason to make that much change when I'm happy enough with Ollama.
- hamdingers 6mo agoWhy not? Also serious. It seems to just work every time I try to use it, the API is easy to work with, the model library is convenient. I've never hit any kind of snag that makes me look elsewhere.
- fennecfoxy 6mo ago
- darshanmakwana 6mo agoReally nice to see this!
- franze 6mo agoI created "apfel" https://github.com/Arthur-Ficial/apfel https://github.com/Arthur-Ficial/apfel a CLI for the apple on-device local foundation model (Apple intelligence) yeah its super limited with its 4k context window and super common false positives guardrails (just ask it to describe a color) ... bit still ... using it in bash scripts that just work without calling home / out or incurring extra costs feels super powerful.
- AbuAssar 6mo agonice project, thanks for sharing. any plans for providing it through brew for easy installation?
- franze 6mo agogood idea
- woadwarrior01 6mo agoThere's a very similar afm CLI that can be installed via Homebrew. https://github.com/scouzi1966/maclocal-api https://github.com/scouzi1966/maclocal-api
- grosswait 6mo agoLooks like they just added homebrew tap to the instructions
- LeoDaVibeci 6mo ago
- techpulselab 6mo ago[flagged]
- harel 6mo agoWhat would be the non Mac computer to run these models locally at the same performance profile? Any similar linux ARM based computers that can reach the same level?
- sgt 6mo agoNot even close. If you want to run this on PC's you need to get a GPU like 5090 but that's still not the same cost per token, and it will be less reliable and use a lot more power. Right now the Apple Silicon machines are the most cost effective per token and per watt.
- harel 6mo agoIt's odd no manufacturer jumped on this wagon to offer a competitive alternative.
- hu3 6mo agoIs there even enough market for this? These models are dumber and slower than API SoTA models and will always be. My time and sanity is much more expensive than insurance against any risk of sending my garbage code to companies worth hundreds of billions of dollars. For most, it's a downgrade to use local models in multiple fronts: total cost of ownership, software maintenance, electricity bill, losing performance on the machine doing the inference, having to deal with more hallucinations/bugs/lower quality code and slower iteration speed.
- zozbot234 6mo ago> These models are dumber and slower than API SoTA models and will always be. Sure but you're paying per-token costs on the SoTA models that are roughly an order of magnitude higher than third-party inference on the locally available models. So when you account for per-token cost, the math skews the other way.
- harel 6mo agoActually yes. For example, I run local models for ingested documents, summaries, etc. The local models are fine, and there is no need for me to pay for tokens. Performance is adequate for that purpose as well. There are many other cases where I run at scale, time is flexible so things can move slower, and I rather keep it all in house. I'm not even getting into areas where data cannot leave the premises for legal reasons. Right now I'm limited with GPUs mostly. But if that world of local models on Apple silicon is so "good", there is room to expand it to other fruits...
- janandonly 6mo ago> Please make sure you have a Mac with more than 32GB of unified memory. Yeah, I can still save money by buying a cheaper device with less RAM and just paying my PPQ.AI or OpenRouter.com fees .
- zozbot234 6mo ago> Please make sure you have a Mac with more than 32GB of unified memory. The lack of proper support for SSD offload (via mmap or otherwise) is really the worst part about this. There's no underlying reason why a 3B-active model shouldn't be able to run, however slowly, on a cheap 8GB MacBook Neo with active weights being streamed in from SSD and cached. (This seems to be in the works for GGML/GGUF as part of upgrading to newer upstream versions; no idea whether MLX inference can also support this easily.)
- daveorzach 6mo agoWhat are significant differences between Ollama and LM Studio now? I haven’t used Ollama because it was missing MLX when I started using LLM GUIs.
- domh 6mo agoI have an M4 Max with 48GB RAM. Anyone have any tips for good local models? Context length? Using the model recommended in the blog post (qwen3.5:35b-a3b-coding-nvfp4) with Ollama 0.19.0 and it can take anywhere between 6-25 seconds for a response (after lots of thinking) from me asking "Hello world". Is this the best that's currently achievable with my hardware or is there something that can be configured to get better results?
- Octoth0rpe 6mo ago> it can take anywhere between 6-25 seconds for a response (after lots of thinking) from me asking "Hello world". That's not an unsurprising result given the pretty ambiguous query, hence all the thinking. Asking "write a simple hello world program in python3" results in a much faster response for me (m4 base w/ 24gb, using qwen3.6:9b).
- hbbio 6mo ago[dead]
- zozbot234 6mo ago> it can take anywhere between 6-25 seconds for a response (after lots of thinking) from me asking "Hello world". Qwen thinking likes to second-guess itself a LOT when faced with simple/vague prompts like that. (I'll answer it this way. Generating output. Wait, I'll answer it that way. Generating output. Wait, I'll answer it this way... lather, rinse, repeat.) I suppose this is their version of "super smart fancy thinking mode". Try something more complex instead.
- drob518 6mo agoIndeed. Qwen doesn’t just second guess itself, it third and fourth guesses itself.
- Kichererbsen 6mo agoSolid Terry Pratchett reference right there.
- androiddrew 6mo agoGet turboquant 4 bit implemented and this would be game changer.
- dev_l1x_be 6mo ago> Please make sure you have a Mac with more than 32GB of unified memory. Time for an upgrade I guess. If I can run Qwen3.5 locally than it is time to switch over to local first LLM usage.
- harrouet 6mo agoAs being on the market for a new mac and comparing refub M4 Max vs M5 _Pro_, I am interested in how much faster the neural engines are -- compared to marketing claims.
- pram 6mo agoM4 Max is going to be faster.
- deleted 6mo ago[deleted]
- obelai 6mo ago[dead]
- deleted 6mo ago[deleted]
- noritaka88 6mo ago[flagged]
- Aurornis 6mo ago> The local inference story is getting real — I've been running 9 autonomous agents on a Mac mini (Haiku via API, not local yet) and the biggest bottleneck isn't the model, it's the coordination layer between agents. Identity, settlement, who-did-what. What does this comment have to do with MLX or the story? Actually, is this just an LLM posting too? This has em-dashes, “it’s not this, it’s that”, and a rule of three statement at the end. EDIT: Account is posting multiple long comments on different threads only 1-2 minutes apart. This is a bot.
- a-dub 6mo agois local llm inference on modern macbook pros comfortable yet? when i played with it a year or so ago, it worked fairly ok but definitely produced uncomfortable levels of heat. (regarding mlx, there were toolkits built on mlx that supported qlora fine tuning and inference, but also produced a bunch of heat)
- Casteil 6mo agoIt's gotten significantly better with the advent of local/offline MoE models (e.g. qwen3.5:35b-a3b, qwen3:30b-a3b, gpt-oss:20b-3.6b), which offer a good balance of prompt response speed and output quality. 'Dense' models of yesteryear (e.g. llama:70b, gemma2/3:27b) tend to be significantly slower by comparison, therefore, your hardware spends a lot more time 'maxed out' for a given prompt.
- abu_ameena 6mo agoOn-device models are the future. Users prefer them. No privacy issues. No dealing with connectivity, tokens, or changes to vendors implementations. I have an app using Foundation Model, and it works great. I only wish I could backport it to pre macOS 26 versions.
- whazor 6mo agoObviously hardware wise the real blocker is memory cost. But there is no reason why future devices couldn't bundle 256GB of mem by default.
- michaelmior 6mo ago> no reason why future devices couldn't bundle 256GB of mem by default Cost is a pretty big reason.
- raw_anon_1111 6mo agoUsers don’t care about “privacy”. If they did, Meta and Alphabet wouldn’t be worth $1T+. Users really don’t matter at all. The revenue for AI companies will be B2B where the user is not the customer - including coding agents. Most people don’t even use computers as their primary “computing device” and most people are buying crappy low end Android phones - no I’m not saying all Android phones are crappy. But that’s what most people are buying with the average selling price of an Android phone being $300.
- barelysapient 6mo agoDifferent users. Many people care about privacy and aren’t using Meta products. And many businesses care about it too and have information policies to protect their IP.
- raw_anon_1111 6mo ago70% of the world’s population use at least one Meta property at least once per day. How many of the other 30% are too poor/young/computer illiterate to be part of an addressable market? Every company has dozens of SaaS products that store their business critical information. Amazon installs Office on each computer, Slack (they were moving away from Chime when I left), and the sales department uses SalesForce - SA’s and Professional Services (former employee). The addressable market of even companies that care about privacy is not a large addressable market. How long will it be before computers become cheap enough that can run even GPT 4 level LLMs that companies will give it to all of their developers?
- ranjeethacker 6mo agoI used today, working nicely.
- braum 6mo agoHow does Ollama help with Claude Code? Claude code runs in terminal but AFAIK connects back to anthropic directly and cannot run locally. I hope I'm missing something obvious.
- EagnaIonat 6mo agoYou can create an MCP to call out to Ollama. Then have Claude farm work out to local models where the raw power isn't required. You can then have Claude review the work from the model. Its not 100% offline, but there is a dramatic drop in token usage. As long as you can put up with the speed.
- navigate8310 6mo agoI believe one can use the CC as the primary model driving local agents that use local models
- samuel 6mo agoYou can connect it to any anthropic compatible endpoint(kimi allows this) but it's a weird choice, given that Open code, pi.dev and others are open source.
- 0xc133 6mo agohttps://docs.ollama.com/integrations/claude-code https://docs.ollama.com/integrations/claude-code You can use models like qwen3.5 running on local hardware in ollama and redirect Claude to use the local ollama API endpoint instead of Anthropic’s servers.
- xmddmx 6mo agoOn a M4 Pro MacBook Pro with 48GB RAM I did this test: ollama run $model "calculate fibonacci numbers in a one-line bash script" --verbose Model PromptEvalRate EvalRate ------------------------------------------------------ qwen3.5:35b-a3b-q4_K_M 6.6 30.0 qwen3.5:35b-a3b-nvfp4 13.2 66.5 qwen3.5:35b-a3b-int4 59.4 84.4 I can't comment on the quality differences (if any) between these three.
- rurban 6mo agoDoes that mean they are now finally a bit faster than llama.cpp? Cannot believe that.
- bwfan123 6mo agoWhat is the cheapest usable local rig for coding ? I dont want fancy agents and such, but something purpose built for coders, and fast-enough for my use, and open-source, so I can tweak it to my liking. Things are moving fast, and I am hesitant to put in 3-4K now in the hope that it would be cheaper if i wait.
- xiphias2 6mo agoIt doesn't look like RAM, CPU GPU or bandwidth is getting cheaper if that helps you, quite the opposite.
- KerrickStaley 6mo agoI think (without having done extensive research) that some sort of Apple hardware is your best bet right now. Apple hasn’t raised RAM upgrade prices [1] (although to be fair their RAM upgrades were hugely inflated before the crunch) and their high memory bandwidth means they do inference faster than most consumer GPUs. I have an M4 MacBook Air with 24 GB RAM and it doesn’t feel sufficient to run a substantial coding model (in addition to all my desktop apps). I’m thinking about upgrading to an M5 MacBook Pro with much more RAM, but I think the capabilities of cloud-hosted models will always run ahead of local models and it might never be that useful to do local inference. In the cloud you can run multiple models in parallel (e.g. to work on different problems in parallel) but locally you only have a fixed amount of memory bandwidth so running multiple model instances in parallel is slower. [1] https://9to5mac.com/2026/03/03/apple-macbook-price-increase-ram-same/ https://9to5mac.com/2026/03/03/apple-macbook-price-increase-...
- victords 6mo agoAs mentioned before, I think Apple hardware is the best alternative right now. Mac Studio, Mac Mini, MacBook Pro, you can find even some used ones with enough RAM that will run models like Qwen reasonably well. I'm using a M1 Max MacBook Pro and it runs Qwen 3.5 on Ollama (without MLX) at a decent speed.
- jwr 6mo agoTwo things: 1) MLX has been available in LM Studio for a long time now, 2) I found that GGUF produced consistently better results in my benchmarking. The difference isn't big, but it's there.
- techpulselab 6mo ago[flagged]
- DevKoan 6mo agoThe Foundation Model point is real. As an iOS developer, what excites me most isn't the performance — it's what on-device inference does to the app architecture. When you're not making network calls, you stop thinking in "loading states" and start thinking in "local state machines." The UX design space opens up completely. Interactions that felt too fast to justify a server round-trip are suddenly viable. The backporting issue is painful though. I've been shipping features wrapped in #available(iOS 26, *) and the fallback UX is basically a different product. It forces you to essentially maintain two app experiences. Still think this is the right direction — especially for junior devs just learning to ship. Fewer moving parts, less infrastructure to debug.
- peronperon 6mo agoDon't post generated comments or AI-edited comments. HN is for conversation between humans. https://news.ycombinator.com/newsguidelines.html#comments https://news.ycombinator.com/newsguidelines.html#comments
- subarctic 6mo agoWhat gave this one away — just the em dashes?
- adolph 6mo agoMuch of the discussion here is local versus remote. I like seeing things as "and" and "or." There will be small things I don't want to burn my Claude tokens on and other things that I want to access larger compute resources. And along the way checking results from both to understand comparative advantage on an ongoing basis.
- jiehong 6mo agoThis is excellent news! What I'm waiting for next is MLX supported speech recognition directly from Ollama. I don’t understand why it should be a separate thing entirely.
- skwon816 6mo ago[dead]
- pyinstallwoes 6mo agoWhat’s the best local coding model these days?