6 ms·
The numbers I was most interested in are tucked away in a chart towards the bottom - the speed comparison of the Mac Studios v.s. a RTX 5090: Qwen3.8 27B tok
by simonw 10d ago
The numbers I was most interested in are tucked away in a chart towards the bottom - the speed comparison of the Mac Studios v.s. a RTX 5090:
Qwen3.8 27B tokens/sec generation speed
Prompt size 8K 64K 128K 256K
RTX 5090 PC 59 51 44 n/a
M5 Ultra 48 39 32 24
M3 Ultra 31 23.5 20 15
A whole bunch more comparison numbers in this section: https://www.macstories.net/stories/m5-ultra-mac-studio-review-the-dream-mac-for-local-ai-agents/#mx-pc https://www.macstories.net/stories/m5-ultra-mac-studio-revie...
- peri-cl 10d agoThose are some incredible graphs, that leap in prompt processing going from M3 to M5. Also: ~30 token/s on GLM 5.3-flash, locally. (That's roughly Opus 4.8-tier. I think). /meta Here's a CSS filter that stops those nuisance chart animations, macstories.net##*:style(animation: none !important; transition: none !important)
- redox99 10d agoA dense 27B doesn't really make sense for the Mac. A MoE makes way more sense when you have modest bandwidth but lots of memory.
- peri-cl 10d agoThey do MoE. They benchmarked GLM 5.3-flash (320B / 18B), and Qwen 3.8-flash-next (125B / 6B). The dense Qwen is only focused (I assume) because it's about the only thing that fits on a 5090, that they can compare the two heads on.
- tcdent 10d agoA dense model (up to the amount of memory available) actually does make the most sense on unified memory architectures. But when you hit the limit of what you can hold in memory, you reach the limitation of the platform. Whereas a hybrid architecture with distinct DRAM and VRAM with sparse MoE, you can leverage two different bit rates depending on the actual need for constant access to common layers versus sparse access to infrequent layers and arbitrage the difference in cost for each of those in distinct classes of hardware.
- gpugreg 10d agoThose RTX 5090 numbers are bad. You can get over 200 tps with ninfer using NVFP4 and MTP.
- beastman82 10d agocan confirm. I dont' know why people spend huge money on these and Spark. The 5090 is running qwen 3.8 at 200+ tps!! That's 1-2 orders of magnitude faster.
- tomega2134 10d agoIs a 5090 still cost efficent when it is (currently) unobtainable? Or when obtainable only at current prices (min. $6500 USD)?
- mathisfun123 10d agosame reason they spend huge amounts of money on rolexes when seikos work better (the tech crowd isn't immune from vanity).
- throwaway27448 10d agoIf you seriously think apple products are nothing but a status item, you're deluding yourself and probably have been for decades.
- deleted 10d ago[deleted]
- _hugerobots_ 10d agoThis 1000%. Data centres don't equate to medium sized labs and businesses. A stack of Macs is up and running without digging trenches, an electrician on staff and a department of PhDs to justify the spend.
- bigyabai 10d ago
- RationPhantoms 10d agoThank you for this. I wish Apple focused their silicon design on improving the TTFT metrics but coming from an M3 Pro, it still looks laggard compared to Nvidia's TensorCores in the 5090. Maybe Apple is an acquisition away from changing that balance.
- wlesieutre 10d agoThe rumor on Apple's processor roadmap is that they're skipping other M6 variations (all previous generations had Pro and Max, a few had Ultra) in order to focus on the M7 generation for AI reasons. What exactly the M7 improvements are who knows. https://www.macrumors.com/2026/06/25/2027-macs-m7-chips/ https://www.macrumors.com/2026/06/25/2027-macs-m7-chips/
- kridsdale1 10d agoI think that comes down to TSMC. Nvidia apparently booked out the whole A18 or 16 node. Apple is on 2nm right now and M7 will jump right to A14. According to my quick AI research anyway.
- dagmx 10d agoThat sounds a lot like AI fantasy slop. Apple just shifted to N2. They’re not going to be doing another major shift right away. And TSMCs own roadmap would put your hallucination years away at best for a a product that follows a roughly annual cadence https://www.tomshardware.com/tech-industry/semiconductors/tsmc-unveils-process-technology-roadmap-through-2029-a12-a13-n2u-announced-a16-slips-to-2027 https://www.tomshardware.com/tech-industry/semiconductors/ts...
- smith7018 10d agoYeah, Apple has spent 3 years each on the 5nm and 3nm nodes with TSMC. There are some reports [1] that it will jump to 1.4nm after 2 years due to AI but there's no real proof. The source is Digitimes who are frequently wrong with their predictions and rumors. [1] https://wccftech.com/apple-to-move-to-1-4nm-process-soon-to-secure-adequate-chip-supply/ https://wccftech.com/apple-to-move-to-1-4nm-process-soon-to-...
- jmyeet 10d agoThe selling point of the M5 Ultra Mac Studio is that you can run much larger models that the 5090 can't without swapping. NVidia aggressively segments the market on VRAM for this reason. That's why a 5090 has an MSRP of ~$2k (but good luck getting one for less than $4k) while a 6000 Pro, which is basically a 5090 with 96GB of RAM has now soared beyond $15k where 3-6 months ago it was more like $10-11k. A 6000 Pro has the same memory bandwidth but slightly more CUDA units (IIRC ~24k vs ~21k). This advantage won't be apparent with a 27B model. The 256GB MS can probably run the newer Flash models locally, something you can't do on a 5090. I don't think we'll get a successor to the 5090 until late 2028, maybe even 2029. I'm basing this on the launch date of the 5000 series and that we haven't got a midcycle refresh yet. Rumor has it the chips are ready but the 3GB RAM modules are 3-4x the price of the 2GB modules used on the current cards. Apple should see a Mac Studio major update in 2028. That might even force NVidia's hand. But it's really impossible to say what the state of the market will be 2-3 years from now. It may have completely crashed. I suspect not however. The interesting thing will be when the bandwidth demands start forcing HBM memory onto these home/enthusiast solutions.
- deleted 10d ago[deleted]
- pama 10d agoBut what about builds that combine 8 of the 5090 with infiniband between boxes? Wouldn't that be comparable to the mac in terms of price and potentially beat it by a lot in terms of performance for the large MoE? I understand the space/heat/noise considerations, but price wise it may still not make as much sense as people think. (Agreed that it is hard to get the NVIDIA hardware and the 6000 pro are priced less competitively).
- kridsdale1 10d agoWhile that sounds super awesome, How many people are actually going to build and maintain that vs a box you can grab at the mall that fits in a lunchbox?
- wmf 10d ago
- traceroute66 10d agoNot forgetting of course that an RTX5090 is what 600W+ ? And the Mac is probably half that at most ?
- beastman82 10d agosure. so is 2x power worth 10x perf? I think it is in most cases.
- washadjeffmad 10d agoCertainly not forgetting wattage. A 5090 is 575W. The M5 Ultra Studio is 480W. nvidia-smi -pl 450 for like a 4% reduction in throughput. I tend to set it around 350W because it's a comfortable temperature blowing on my legs under the desk without warming my office in the summer. I put together this system two years ago, so it's a little out of date, but it only cost $3000 for the same performance and capability as an Ultra. I don't think I would spend $7000 to save 100W, though.
- TacticalCoder 10d ago> nvidia-smi -pl 450 for like a 4% reduction in throughput. Yeah people don't pay enough attention to those settings IMO. The first thing I do when I set up a new machine (or upgrade my OS) is to restore all my powersaving configs. For example I've got all but one of my virtual desktops that put the CPU in powersave mode: I don't need max Ghz when browsing the Web, not even on demand. But when I switch to the virtual desktop where my development environment is, then I want power on demand. Now I don't do it to save the planet: I do it because I love a quieter computing experience (coupled with Be Quiet! PSU and Noctua fans, this makes for a very quiet computer). That it consumes less electricity is a nice side-benefit.
- ActorNightly 10d agoWhen you are doing matrix math, compute is compute. Apple cant be more efficient due to physics. The only reason Macs are more efficient in general is that they have tightly bundled hw and sw for specific tasks.
- nacs 10d agoThat's a dense model. Of course it will do worse. Now try running that Qwen 3.8 Next model on the 5090 and tell me what TPS you get (hint: it's near 0 since it doesnt fit the 32GB VRAM on 5090 vs the 256 in OPs M5).
- peri-cl 10d agoSurprisingly, the Reddit crowd are reporting 50–60 tokens/s (for the 32 GiB 5090 + 128 GiB RAM)—on par with the M5 Ultra benchmarks, despite both the PCIe bottleneck and much smaller DDR5 bandwidth, https://old.reddit.com/r/LocalLLaMA/comments/1wl06np/qwen38flashnext_on_1x_rtx_5090_tg50_ts_pp2300_ts/ https://old.reddit.com/r/LocalLLaMA/comments/1wl06np/qwen38f... (Note it's a sparse MoE with only 6B active).
- nacs 10d agoGood to know thanks. That's with CPU offload to a DDR5 6000 RAM though which is around $3-4k at least.
- well_ackshually 10d agoUnlike a 256GB M5 Ultra that is $10k+.
- nacs 10d agoApple product won't be the cheapest but it is a full package (CPU, RAM, VRAM/GPU, fast-storage, etc). If you look at the pricing of a full (x86) AI workstation you'd need around the nvidia GPU, you'd approach $10k easily (and be using a ton more wattage too).
- api 10d agoI assume those are non-batched. I think the M series GPU can do 4X to 8X depending on model quant, which means if you can batch queries you'll get almost 4X to 8X performance.
- alex7o 10d agoOn my m5 max 27b model does 75tps on 256k ctx and starts at 80 on the 8k ctx when you add https://huggingface.co/collections/z-lab/dflash-2 https://huggingface.co/collections/z-lab/dflash-2 to it. So yeah base might be 30tps (I used iq4) but mtp or dflash help a lot and should be used when checking what is useful and what is not for running models as it is not fare to judge without them.
- GeekyBear 10d agoThe next Ultra, supposedly on deck in 2028: > Apple's planned M7 Ultra chip is being designed to support up to 1.5 TB of unified memory and to push AI performance toward the class of Nvidia's Blackwell accelerators, according to a new Bloomberg report published by Mark Gurman... Apple plans to release a base M6 chip this fall for entry-level Macs... a base M7 in the first half of 2027, M7 Pro and M7 Max at the end of 2027, and the M7 Ultra in 2028. https://www.tomshardware.com/tech-industry/semiconductors/apples-rumored-m7-ultra-targets-1-5tb-of-memory-and-blackwell-class-ai https://www.tomshardware.com/tech-industry/semiconductors/ap...
- karmakaze 10d agoI really appreciate seeing these dense model numbers. For a large unified memory system though I expect that MoE numbers are what people are more interested in. These numbers could and should get much better. As an example I can run Qwen3.8-27B-MXFP4 (W4A8) on 2x AMD R9700 that gets 260+ tokens/sec to start and slows down to ~110 tokens/sec over 128k context and can do the max 256k. These are for batch size 1 and throughput goes higher with batching. This is due to speculative decoding, efficient all-reduce inter-gpu compression, and custom GEMM kernels for the specific hardware. Note each R9700 only has 644 GB/s memory bandwidth.