5 ms·
M5 Ultra Mac Studio Review
- saejox 15d agoi can buy a house with that amount of money. it used to be car money.
- ApolloFortyNine 15d agoThe model being tested is 18k as configured. I didn't expect this to make the 5090 to look like a good deal.
- nacs 15d ago5090 has 32GB VRAM. It'd be silly to buy the 18k model to run a tiny model like Qwen 27B. You use models like GLM Flash and Qwen Next which won't fit on a single 5090.
- orsorna 15d agoIs it that silly? You could run multiple 27B models in parallel.
- peri-cl 15d agoYou actually don't need more RAM to batch multiple inference tasks of the same model. (Each task needs its own context, but the (e.g.) 27B of constant parameters isn't duplicated).
- orsorna 15d agoYou definitely need more RAM if you are not satisfied with small context windows, especially if the weights take a large % of the total memory to boot.
- deleted 15d ago[deleted]
- asimovDev 15d agocan run multiple subagents of Qwen 27B though, right? Unless I am fundamentally misunderstanding how VRAM constraints work
- Eisenstein 15d agoYou might be. Running another agent doesn't load a set of new weights. It creates a new KV cache for the agent and adds the prompts to the queue. Its just another inference turn.
- rwissinger 15d ago[dead]
- snarfy 15d ago$12,299
- andrekandre 15d ago5 years of (200/month) tokens at that price, meanwhile an rtx 5090 pc is about half that… hmm but i wonder how much these token costs are sustainable or not, it may be in the long term cheaper to have your own hardware if token costs go up (and hopefully hardware gets cheaper again)
- chasd00 15d agoThe token price isn't the only reason to run a model locally though. You can do additional training to specialize or remove censorship that may be a no-no per TOS with cloud GPUs.
- drdaeman 15d agoCloud GPUs have ToS about purposes you’re allowed to crank numbers for? I thought this only applies to LLM inference providers, but not raw GPU rentals.
- Eisenstein 15d ago$2200 for a 64GB VRAM machine if you are willing to do a bit of work. * https://imgur.com/mpdorVJ https://imgur.com/mpdorVJ
- drdaeman 15d agoThose tokens aren’t guaranteed (esp. with RE and security tasks - rooted my own TV last week, Claude crapped out on “cyber safety” grounds), and you’re throwing money at entities that aren’t aligned with your interests instead of entities who are interested in actually empowering you.
- simonw 15d agoThe numbers I was most interested in are tucked away in a chart towards the bottom - the speed comparison of the Mac Studios v.s. a RTX 5090: Qwen3.8 27B tokens/sec generation speed Prompt size 8K 64K 128K 256K RTX 5090 PC 59 51 44 n/a M5 Ultra 48 39 32 24 M3 Ultra 31 23.5 20 15 A whole bunch more comparison numbers in this section: https://www.macstories.net/stories/m5-ultra-mac-studio-review-the-dream-mac-for-local-ai-agents/#mx-pc https://www.macstories.net/stories/m5-ultra-mac-studio-revie...
- peri-cl 15d agoThose are some incredible graphs, that leap in prompt processing going from M3 to M5. Also: ~30 token/s on GLM 5.3-flash, locally. (That's roughly Opus 4.8-tier. I think). /meta Here's a CSS filter that stops those nuisance chart animations, macstories.net##*:style(animation: none !important; transition: none !important)
- redox99 15d agoA dense 27B doesn't really make sense for the Mac. A MoE makes way more sense when you have modest bandwidth but lots of memory.
- peri-cl 15d agoThey do MoE. They benchmarked GLM 5.3-flash (320B / 18B), and Qwen 3.8-flash-next (125B / 6B). The dense Qwen is only focused (I assume) because it's about the only thing that fits on a 5090, that they can compare the two heads on.
- tcdent 15d agoA dense model (up to the amount of memory available) actually does make the most sense on unified memory architectures. But when you hit the limit of what you can hold in memory, you reach the limitation of the platform. Whereas a hybrid architecture with distinct DRAM and VRAM with sparse MoE, you can leverage two different bit rates depending on the actual need for constant access to common layers versus sparse access to infrequent layers and arbitrage the difference in cost for each of those in distinct classes of hardware.
- sajithdilshan 15d agoOn Apple website it says 512GB memory option is available in October. I guess bumping to that one would cost additional 4-6k US$. So an Ultra with 2TB storage would be north of 15k US$. That’s like 12 years worth of OpenAI Pro subscriptions
- geodel 15d agoAgreed. Specially since one can pay half right now to OpenAI and sign a 12 year iron clad contract for uninterrupted service delivery of OpenAI Pro.
- vardump 15d agoI hope that was sarcasm.
- cyclopeanutopia 15d ago"iron clad" :)
- prmoustache 15d agoCensorship included.
- Kurtz79 15d agoI think we all expect the heavy subsidized subscriptions to end or significantly increase in price at some point, but it could be years from now and I'd rather spend a similar figure on an hypotetical Mac Studio M8 Ultra, or whatever more advanced competitor that will have likley appeared by that time. A more apples-to-apples comparison would be with API cost in OpenRouter at the same tok/s rate for the same models that you can run locally, maybe.
- BatFastard 15d ago>A more apples-to-apples comparison Don't you mean an Apple to NVidea comparison?
- kokonokko1337 15d ago> "It also happens to be a Mac, with an operating system that looks nice and doesn’t suck" Yes Apple has some of the best hardware out there, albeit overpriced. But the software is such a hindrance and I can't take anyone that states otherwise seriously. If only it had proper Linux support (and the Asahi people do an amazing job but you can reverse-engineer only so many stuff with limited funding, and then you have to do it again for new models). MacOS is good if you just want to have a standard experience, which to be fair is most people. It's good for just setting up an LLM server I guess since the hardware is a perfect fit. I wouldn't touch it otherwise.
- steve1977 15d agoWhat exactly is missing from macOS that makes you feel the need for Linux? I get it on Windows systems, at least when someone wants to use Linux-type tooling. But macOS already supports pretty much all of that natively?
- RunSet 15d ago> What exactly is missing from macOS that makes you feel the need for Linux? For starters, the source code.
- steve1977 15d agoAnd why would you need that to run LLMs? Apart from that, for the UNIX part, the source is available for quite a few components: https://github.com/apple-oss-distributions https://github.com/apple-oss-distributions notably also the kernel https://github.com/apple-oss-distributions/xnu https://github.com/apple-oss-distributions/xnu
- throw0101c 15d ago> Apart from that, for the UNIX part, the source is available for quite a few components: Strictly speaking, Apple can claim to ship a UNIX® operating system: * https://www.opengroup.org/openbrand/register/apple.htm https://www.opengroup.org/openbrand/register/apple.htm * https://www.opengroup.org/openbrand/register/ https://www.opengroup.org/openbrand/register/
- srcreigh 15d agoThis is great as a first look, but the author is not a developer, so we don't yet know whether a dev can be as productive with local models on M5 Mac Studio compared to a 20x subscription plan. I'm also curious about any new low hanging optimization opportunities in the kernels for this new hardware. It's already clear to me that M5 Mac Studio is more cost-effective than anything you can run on open router, assuming decent utilization. The M5 Mac Studio will be the most cost effective way to run uncensored cyber capable open agents. An exciting tipping point will be if programmers can get an Astra-Ultra like experience all week with this hardware. That would be a real sense where this hardware exceeds the value of even 20x cloud subscriptions.
- slowin 15d ago> This is great as a first look, but the author is not a developer, so we don't yet know whether a dev can be as productive with local models on M5 Mac Studio compared to a 20x subscription plan. Local models are definitely not as productive as SOTA, sadly it's not close yet. I do think someday they will be "good enough" to use, but they aren't today. Even the SOTA models barely code well, with Opus 4.5 being the first, good coding model. That being said, I think it's absolutely imperative that we keep pushing local model performance. We need to continue to advance technology there and ensure that the model labs don't do regulatory capture in the name of "safety" (or anything else).
- nowittyusername 15d agoWith the latest codex (weekly quota burn) fiasco I tried open weight alternatives for the first time. And tyeah... open weight models cant compete with likes of astra yet. But, my hope is that by the time I get my Mac studio at end of november an open weight models would have closed the gap (which i think is realistic at the speed of progress). Now its true a better gpt version will also be available then but it also seems the gap is shrinking with time so theres that.
- Octoth0rpe 15d ago> And tyeah... open weight models cant compete with likes of astra yet I think this is true, but also misses that a lot of us are just doing basic flask apps with a react front end. We don't need astra; Something sonnet 4.6 level locally is perfectly sufficient 95% of the time, and maybe 99% of the time.
- WarmWash 15d ago>Let’s address the elephant in the room first: why bother with local AI at all when cloud frontier models are better and often faster? Ehh, the actual elephant in the room is: "why bother with local AI at all when you can lease a GPU for $5/hr?" To which the answer is you shouldn't bother, unless you have a bunch of money to throw at hobby projects.
- chasd00 15d ago> unless you have a bunch of money to throw at hobby projects. there are lots of people with very expensive hobbies, see sailboat racing for example.
- deleted 15d ago[deleted]
- Youden 15d ago$5/hr = $3600/mo. Unless you only need the AI available some of the time, $5/hr is pretty expensive. That's an RTX Pro twice a year. If you're using it for discrete sessions of coding or something, that might make sense for you but if you're using it for an always-on assistant, that pricing kinda sucks.
- WarmWash 15d agoI would imagine extremely few people are utilizing an H200 for every hour of a month. Especially for something like an assistant 5090's are like $0.20/hr
- tempoponet 15d agoWhile I know it's not apples to apples, the target comparison right now is 2x DGX Sparks. Similar price, 256gb. The conversation has focused on memory bandwidth vs. compute in agentic loops, so for most people the raw numbers will mean less than the "time per task" in coding benchmarks. This is a great article and bodes well for the M5, but we should expect more like this comparing to other platforms before we truly understand where it fits.
- _hugerobots_ 15d agoSpeed vs task-completion is a new conversation and a great point. Whereas the cost to compute doesn't exist in a vacuum, making mistakes costs less, is easier to maintain with granularity and a whole host of other factors when you own the lab.
- sghiassy 15d agoImagine spending a trillion dollars on data centers and then reading this article. Nightmare fuel for OpenAI
- whalesalad 15d agoFor 99.99% of people, spending 15 grand on a Mac Studio just to run Qwen 3.8 locally is a non starter.
- jmull 15d agoIt's not the M5 Ultra itself, but the M7s or M9s that will do the damage. 99% of people will use whatever AI is free. The sophisticated, heavy users that are willing and able to pay a lot of money the ones that will be interested in controlling their inference bills. Today, the sweet spot where an M5 Ultra makes sense is tiny. But we might expect that to grow a lot.
- BatFastard 15d agoAnthropic is reporting 100 Billion ARR. Even if you could get a frontier model, you would not be able to run it on any Mac. So speculating on what M7 or M9 will achieve in 5 years (if we even still exist) seems pointless.
- sghiassy 15d agoDo you need a frontier model to write emails, check your calendar, search the web? I don’t think Apple is going to lie down and cede AI to the cloud.
- geodel 15d ago> Do you need a frontier model to write emails, check your calendar, search the web? How about writing mail to President and senators on AI doomsday scenario if frontier labs do not pace themselves? That mini model on mac mini would scared to hell to do such thing. It need that rugged frontier model to speak truth to power.
- devy 15d agoThis dream machine costs over $15k (not including the Apple Studio Display)? Nah, that dream is SO OUT OF TOUCH!
- aenis 15d agoThe irony here is, thats hobby hardware. You spend 20k and can run slow hobby models that are barely capable of anything unsupervised. Entry level serious hardware starts at 100k, and a bit better but still almost-useful grade is 200k (8x rtx pro, plus a nice epyc pairing). Thats the sort of thing a salaried expert lets their employer buy them for sort of serious work. Anything really serious is well north of 1M - not including the housing and commercial grade mains connection. And at best that buys fast Kimi K3 or GLM.
- BatchJob 15d agoIf you are buying expensive hardware to run LLMs "on your own machine" you will soon find your ladder is on the wrong wall.
- addaon 15d agoOrdered one for OpenFOAM. Excited for it. Will be nice to not have my laptop running CFD 24 hours a day, but my M1 Max is currently my fastest machine… I’m expecting about 3.5x from the M5 Ultra.
- SamuelAdams 15d agoI think Apple is really sleeping on making this run a Linux server. These things are very capable and draw very little wattage when idle. It would make an excellent homelab device, but MacOS currently holds it back in this regard.
- flounder3 15d agohttps://mac.getutm.app/ https://mac.getutm.app/
- deleted 15d ago[deleted]
- jjtheblunt 15d agoi use linux a ton too, but still wonder what you want in a Linux server that macos as a BSD server does not have.
- Gracana 15d agoI'll probably manage with Mac OS well enough, but my linux distro comes out of the box with all the latest OSS tooling I'm familiar with, plus a package manager, and it has linux cgroups and namespaces that power the container technologies we all know and love. If I switch to Mac OS, I have to sort out a package manager and install all the stuff that's missing, and when it comes to containers... they're just linux VMs. I'd happily cut out the weird proprietary middleman if I could.
- crossroadsguy 15d agoMy mac is 5 years old. I don't think I can comfortably buy a new one right now. It has a 16GB unified RAM. Honestly that would be enough for so many local models that I want to use but can't use. Because RAM usage (even with literally every single user installed app quit/stopped) the RAM usage is very high that I can barely safely get 6-7 GB (I am supposed to get ~10 GB, but it goes up and down real fast!). That's a shame. If only I could install an alternative OS that uses very little amount of RAM :-)
- mjlee 15d agoHow are you measuring memory usage? top tells me that 45/48GB is "used", but Activity Monitor shows me that 24GB is cached files. I'd be quite surprised if Mac OS alone needs more than 8GB, given that they sell the Neo with 8GB of RAM today.
- liuliu 15d agoWhen people benchmark MLX related quant models, they really need to publish numbers on benchmarks. You cannot take this as it is what you get of the original models. MLX uses pretty simple quantization methods so at lower bits without QAT, it is just not as good quality as llama.cpp ones.
- theplumber 15d agoAt this point I think I will get the DGX gb300 workstation though I will wait a bit more for the cold season. It is double the price but at least is the real thing
- 12kaj2 15d agoThe Year Of Local AI will be here no later than 2040, coinciding with the Year Of The Linux Desktop.
- prmoustache 15d agoThe year of linux on the Desktop was 26 years ago for me.
- akozak 15d ago"a total cost of $0" Uhh ... how much is that hardware?
- novaleaf 15d agoAnother comment approximates at around USD$15k, so yeah, not zero.
- villgax 15d agoLol, try generation of images & videos on these, they ought to improve perf on Deep learning not just llms
- crorella 15d agoWhat are good options to run local models nowadays? Something good for coding and personal assistant kind of things
- cptskippy 15d agoI think we'll eventually get to the point where folks will have a local AI agent but I think people need to temper their expectations to a degree. You aren't going to have data center level tok/s from a box sitting under your desk and you don't need instantaneous responses for many workloads. Having a local agent that can execute tasks over a couple days with your supervision that might otherwise take you weeks is perfectly acceptable. However I also think that Agentic AI is very much not an out-of-the-box solution, local or otherwise, and it takes a high level of technical knowledge to create an effective AI agent. And there's a problem now where most orchestration is fixed on what models are used for what tasks with no ability to weight constraints like cost, speed, and security.
- slashtom 15d agoFantastic review, this is how it should be done with local AI.
- hamiltont 15d agoOnce you hit the memory you need, generation speed is mainly set by bandwidth, and every Ultra from M1 thru M3 has ~800 GB/s. IMO best ROI for most people is 'cheapest used Ultra with enough RAM' I setup an eBay alert and picked up a used M2 Ultra that has delivered good ROI (at least, far better than 15k for comparable-for-my-use-case performance)
- Lwerewolf 15d agoThis one is 2x m5 max, so ~1.2TB/sec.
- peri-cl 15d agoI think M1 through M3 were compute bottlenecked in prompt processing (hence the very large gap between M3 and M5, in this page's benchmarks, that's not explained by memory bandwidth alone). For generation speed in isolation, yes.
- GeekyBear 15d agoThe M5 generation added tensor instructions to the GPU cores.
- itsmeduncan 15d ago[flagged]