6 ms·
Nativ: Run frontier open models locally on your Mac
- rvz 3mo agoLooks like my call [0] for more competitors to Ollama has been answered. We need more like this as well as llama.app, which also has a native mac app. [0] https://news.ycombinator.com/item?id=48968898 https://news.ycombinator.com/item?id=48968898
- laughingcurve 3mo agoAgreed. Just based on this not being Ollama, so I will give it a try.
- 0gs 3mo agomy thing is kind of an Ollama competitor (surrogate?) too. more for prose/text planning, not so much for coding, at least the harness, but i'm sure someone could set it up to do that: github.com/0gsd/enough
- iAMkenough 3mo agoHow does this compare to LM Studio Bionic?
- asqueella 3mo agoBionic is an agent; this appears to be an open-source competitor to lmstudio. The initial commit is just a few hours ago though, so …
- calumcl 3mo agoThe maintainer works on mlx-vlm so he does have pedigree in the scene, I'm don't know if this is just going to end up as unmaintained slop. I haven't tried this but I would personally just recommend oMLX for a currently more complete and fleshed out package - loads of features, provides the same MLX support and changelogs + commits are actually detailed.
- jdiff 3mo agoI use oMLX and I'm tentatively going to be giving this a shot. oMLX keeps driving me up a wall with odd papercuts, bugs, and silent failures and fallbacks that are only visible buried deep inside logs when they should be announced out loud. The maintainer of mlx-vlm being behind this as well is the main thing kicking me over into trying it, even if it is incredibly young. I'm confused and unenthused to see it chomping on a whole GB of disk, but the Swift makes it feel much more refined even if it's not yet as feature rich. It automatically picked up the existing models I was using with mlx_vm directly, which was nifty.
- crefiz 3mo agoWhy do ppl downvote this comment? HN makes no sense sometimes...
- SafeFatNoob 3mo agoso now we have this, https://pypi.org/project/rapid-mlx/ https://pypi.org/project/rapid-mlx/, https://mtplx.com https://mtplx.com, and the oldest I could find at https://omlx.ai https://omlx.ai.
- philips 3mo agoThere is also https://github.com/ml-explore/mlx-swift-lm https://github.com/ml-explore/mlx-swift-lm which is what https://jan.ai https://jan.ai uses.
- moontear 3mo agoIs „frontier“ overused? I thought frontier models were the best-of-the-best such as Fable right now. I assume you can’t host these models yourself since you would need many GB of RAM and expensive GPU of is my thinking of „frontier models“ wrong?
- 44za12 3mo ago+1 came here to say this, I opened the link expecting some technical breakthrough. Misleading click bait title.
- bnfcl 3mo agoHad the same though. The gap between open-source and the frontier is closing in, especially with Kimi K3, but that is like >2T parameters. The Gemma 4 and other models you can actually run on an average Mac, is not in the same league.
- zeckalpha 3mo agoThe frontier is a curve. https://en.wikipedia.org/wiki/Pareto_front https://en.wikipedia.org/wiki/Pareto_front
- IshKebab 3mo agoThat's not what people are normally referring to when they say "frontier models". It means the most capable models full stop. Not the most capable that you can run locally.
- chrisweekly 3mo agotrue but tfa's title says "frontier open models"
- wmf 3mo agoFrontier open models are Kimi K3, GLM 5.2, DeepSeek V4 Pro, etc. They're all too big to fit on most Macs.
- brcmthrowaway 3mo ago[dead]
- dlandis 3mo agoAdvice: remove all slop and fluff from the website such as "Everything you need. Nothing you don’t." Just state the information you want to communicate in the plainest and most straightforward way possible.
- vitally3643 3mo agoSomething nobody needs is pointless hot air like "Everything you need. Nothing you don't."
- thejazzman 3mo agoYou’d be surprised how hard this actually is. I spent 3 days iterating on a marketing site, where I had very explicit / “well written” copy, and it would just repeatedly rewrite it back to the most awful slop. Over and over again! Ended up adding various AGENTS rules telling it to leave the copy alone
- jmpz 3mo agoDefinitely easier and better than just writing the copy..
- lantry 3mo agoI think they're saying that they _had_ written the copy themselves, but were using the clanker for other tasks, and it kept going off course to "improve" the copy.
- satvikpendem 3mo agoJust...write it yourself.
- thejazzman 3mo agoI did. And then it rewrote it. Over and over again. And I kept restoring it. And it kept taking agency to change it to some other neutral slop. That’s the point.
- isomorphic 3mo agoThis looks like the Prism folks, who are making binary/ternary versions of popular edge models, so that those models will fit on constrained devices like phones. E.g., their Bonsai model derived from Qwen: https://news.ycombinator.com/item?id=48910545 https://news.ycombinator.com/item?id=48910545 Perhaps they got tired of LM Studio, etc., not being able to run their models properly.
- dofm 3mo agoIs it? The developer page for the github repo suggests he works at Arcee. Maybe he moved? (I initially thought the same because of the website appearance)
- D13Fd 3mo agoI'm surprised that their home page basically acts as if LM Studio and others don't already do this. It's not clear what the difference is from a glance. It also omits Open WebUI. I've been running Deepseek V4 Flash locally on my Macbook Pro for weeks using Open WebUI + DS4.
- lylejohnson 3mo agoI was wondering the same and assuming I'd overlooked something.
- kzrdude 3mo agoLM Studio seems to do the same thing, except that LM Studio is not open source. So they have a point, they do something more.
- moostii 3mo agoLM studio is closed source software built ON TOP OF code released by the author of Nativ.
- a3w 3mo agojan.ai is the F/LOSS alternative to it. If you do want a gui.
- wmf 3mo ago"The other “local AI” apps you’ve heard of? They’re proprietary shells built on top of open-source engines they don’t own." This is a roundabout way of addressing LM Studio.
- syntaxing 3mo agoWhat spec is your macbook? I want to run Deepseek V4 Flash but its too slow for agents on my Strix Halo.
- D13Fd 3mo agoIt’s a maxed out current-gen MacBook Pro w/ 128gb ram. I lucked out in getting it before they upped the price. It runs really well and fast enough for what I need, and it works great, but it’s slower than Claude. I use it for projects where I’m not allowed to use third party AI for legal reasons.
- deleted 3mo ago[deleted]
- satvikpendem 3mo agoOnly interesting thing about this vibe coded runner is the MLX support, as that's still annoying to use in other ones, most still use GGUFs. Unsloth Studio which is an OSS runner I use is still in progress with MLX support although it's still a ways away.
- audioh4cker 3mo ago[flagged]
- sajithdilshan 3mo agoHas anyone found a model that can run on a normal macbook? I have an M3 Pro with 18GB of memory and whenever I try to run even a basic model the fans goes off and the mac starts to get heated up and becomes so laggy.
- sixtyj 3mo agoTbh, the only reason to run a model locally is when you want to be completely safe. From productivity point of view, it doesn’t make sense to have any notebook running a local LLM. We have one life. We should spend it wisely.
- mft_ 3mo agoThere is a Gemma 4 model with 12B parameters which might be worth trying. e.g. https://huggingface.co/mlx-community/gemma-4-12B-it-qat-4bit https://huggingface.co/mlx-community/gemma-4-12B-it-qat-4bit That said, your computer will still get hot!
- Daunk 3mo agoI just run Gemma4 12B MLX via Ollama and it's been doing fantastic work!
- kzrdude 3mo agoIf you go small enough it should be no problem. For example Gemma 4 E4B in Q6 or Q4 quantization should run well on your laptop. It shouldn't be too taxing, but would still want to eat 7-9 GB of VRAM or so. Now that model is mostly useful for writing or chatting.
- fl0id 3mo agoheating up is normal, that cannot be avoided. it should become laggy, but you just have very little RAM (I assume 16 GB?) So most models are too big with other stuff running.
- b3ing 3mo agoYou need more RAM, plus the models take up a lot of space. 32gb min but I’d recommend 48/64gb, you won’t get close to frontier but it’s still fun to play with, images are very good
- c4pt0r 3mo agoI'm a bit curious why not running DeepSeek V4 on top of https://github.com/antirez/ds4 https://github.com/antirez/ds4. I think the results could be really good.
- Archit3ch 3mo agoI assume Nativ doesn't support SSD streaming like DwarfStar.
- erikkahler 3mo agoTo clarify, this MIT-licensed app is from the very same dev, 'Prince Canuma', who maintains the popular MLX-VLM library (https://github.com/Blaizzy/mlx-vlm https://github.com/Blaizzy/mlx-vlm). MLX-VLM is a long-time dependency of the excellent LM Studio and others because it can provide faster inference on Apple devices than llama.cpp. Historically, MLX is a smaller community than CUDA, but has some of the fastest updates upon the release of new models, particularly in models with modalities beyond text-in, text-out (vision, STT, TTS, video gen). See also (https://github.com/Blaizzy/mlx-audio-swift https://github.com/Blaizzy/mlx-audio-swift). Would be totally unsurprised if those modalities and models get integrated into this UI. Possibly vibe-coded landing page notwithstanding, the app is mostly written in Swift language. That suggests it will be easy to port this inference stack to iPad and iPhone.
- trollbridge 3mo agoThanks for that. My first question was “What does this do that Unsloth doesn’t?”
- recursivegirth 3mo agoLM Studio is trash on Windows / Linux... guess that makes sense...
- Forgeties79 3mo agoWhat makes it trash on Linux? I’m a pretty casual user myself so my guess is I haven’t bumped against these limitations. I also don’t have high expectations as I’m running it on a 9070 lol
- rahimnathwani 3mo agoWould be totally unsurprised if those modalities and models get integrated into this UI. Yup, the GitHub repo says: Support for dedicated audio-only and image-generation-only models is coming soon. Prince Canuma is super-responsive on X and GitHub issues, and I use mlx-audio almost daily with mlx-community/Qwen3-TTS-12Hz-1.7B-Base-bf16 (for voice cloning).
- 3mo ago
- jarek83 3mo agoOk, so we're at the point that even design are just one-shot by AI. This is exact look and feel any time I ask it to present a HTML doc about anything.
- moostii 3mo agoThe webpage isn't the point. The software the webpage refers to is the point.
- woadwarrior01 3mo agoIronic that the app is named Nativ(e) and yet bundles a full Python runtime. Nonetheless, still less bloated than LM Studio, which bundles a full Python runtime and electron.js (which in turn bundles a whole browser runtime).
- bigyabai 3mo agoA lot of macOS apps are statically-linked, even interpreted programs. It's still a native app for going that route.
- Daunk 3mo agoIs Gemma 4 E2B actually "usable"? I've been running Gemma 4 12B and it handles everything very well! But the second I've moved down to E4B it's been unable to perform the simplest of tasks. So I can't even imagine how E2B would do... Or am I doing something wrong?
- JosNun 3mo agoGenuinely curious: what are people using these smaller local models for? They are getting decently capable, but they are still small enough that I don't trust them for "real" work outside of a handful of fun toy projects. Are people actually using them in coding agents? Or are they mostly using them for other things?
- teaearlgraycold 3mo agoThey're great at helping me look up web dev stuff when I don't have internet access.
- kgeist 3mo agoWe've shipped some code generated by Qwen3.6 27B to production (under OpenCode). It lacks the breadth of knowledge of models like Opus, but if a change is fully inferable from the prompt and the surrounding code, it works very well. It won't be able to write something from scratch that requires niche knowledge (say, a performant inference engine tailored to Blackwell GPUs), but if it's just a PR adding a new use case to an existing project (which is usually just "load from the DB, do some invariant checks, modify the entities, store them back"), it works as well as Sonnet (provided you have the correct configuration, like recommended temperature and top-p settings, the model isn't over-quantized, you have at least 150k tokens of context available, etc.).
- jiqiren 3mo agothere is plenty of grunt work these smaller models can do. update dependencies, fix merge conflicts, write --help, markdown, or readme files for existing code. etc. sometimes they fail but undo is just a "git restore" or if automated, rejecting a PR and having a better model take a crack at it.
- mips_avatar 3mo agoQwen35ba3b can do a huge amount of data cleaning work on pretty modest hardware. Already have run about 100 billion tokens on it using 2x3090 gpus.
- febed 3mo ago
- shitcoder 3mo agoLooking forward to giving this a try. I have tried MLX using Rapid MLX however the LLM (Qwen) would always have hiccups and get stuck repeating itself. Moving onto llama.cpp I was able to get faster tokens with MTP and a more reliable llm. I wonder what other people's experiences are using MLX vs llama.cpp
- dofm 3mo agoFWIW on my M1 Max I have not really seen any advantage at all from MLX. I am fully prepared to believe the benefits accrue more to the M3 and up (because of changes to the Apple Neural Engine). But with the models I've tested, unless I am missing something, the performance of GGUFs in llama.cpp has been better in some cases. I still have not had results from Gemma 4's MTP be really worth it, to be honest; but with the Qwen 3.6 MoE it is measurable. Maybe with newer kit it is more meaningful. (There is every chance that the above is not the experience of anyone who really deeply knows what they are doing; it feels like I am a perpetual novice at this stuff)
- regexorcist 3mo agoSame here. Tried MLX twice at different times after reading the claims here but it always does considerably worse for me than llamacpp.
- deleted 3mo ago[deleted]
- VaporJournalAPP 3mo ago[flagged]
- TechSquidTV 3mo agoThe server wont start for me. "ERROR: Application startup failed. Exiting." "mlx-vlm-server stopped with status 3"
- troygentic 3mo ago[flagged]
- jsomedon 3mo agoI can't really tell why would I use this over like lm studio, jan.ai and such?
- flyingcapabara 3mo ago[dead]
- flyingcapabara 3mo agoIf you are on Linux check out Box, it runs models locally , has img gen etc and many more features, I will be releasing the source soon just polishing out last bug's, help porting to other distributions is welcomed Github.com/jegly/b0x
- Atlas_Smith 3mo ago[flagged]
- troygentic 3mo ago[flagged]
- saagarjha 3mo agoI'm a little confused why the page lists "UNIVERSAL · APPLE SILICON (M1+)". This seems like an oxymoron
- khurs 3mo agoFlagged this, as it is not 'frontier models' in the title of the linked page so not sure why has been titled that way here. It's a further way to run mlx models mlx are apple specific format for M3 or later CPUs, and some benchmarks show mlx are not always better than just running generic ones.
- noja 3mo agoCan you add support to download gated models?
- shireboy 3mo agoWhat is the “middlest” Mac one could get for this? I’m in the market but keep going back and forth between a 64gb m5 pro or “lower end m5 air and screw it I’ll just pay for cloud tokens”. At current prices the 2-3k diff to try to run something local that isn’t as powerful could buy a lot of tokens.
- froobius 3mo agoYou can run a lot of these on e.g. M1 max 64gb
- Nekorosu 3mo agoI really don't like the marketing texts. "Why we’re open source when nobody else is." I'm using oMLX which is open source and seems to be doing everything Nativ offers. I'd rather see the comparison with existing "non-existing" open source competitors.
- giancarlostoro 3mo agoWhat they likely mean is, why options like LM Studio are not open source.
- kmike84 3mo agoomlx is quite similar to LM Studio, so there are "options like LM Studio" which are open source
- diimdeep 3mo agoI am still rocking Sequoia and this targets Tahoe purely from UI constraints, no like.
- cootsnuck 3mo agoSo should I ditch ollama for this?
- mfro 3mo agoI don't love that this starts an API server on launch and has no option to disable it...
- n8henrie 3mo agoAny reason to use this over omlx?
- reagle 3mo agoThat's what I wondered.
- jwr 3mo agoHow is it better than, say, LM Studio? And does it run MTP models (which from what I understand are GGUF)?
- themihai 3mo agoWhat’s the BS with “Open models from teams we trust.”?
- m3kw9 3mo agoThe site design itself is likely one shotted. The italic fonts is nauseating
- jwr 3mo agoThe word "frontier" is like "load-bearing" at this point (Claude Code users will know what I mean). I wish we could stop using it. Especially as this does not, in fact, run the leading/top models locally on your Mac.
- przemarzec 3mo ago[dead]
- it 3mo agoI ran it with Qwen3-8B, and it responded to all input in the chat with Could not connect to the server.