4 ms·
I was already rolling around the idea of a 128GB M5 Max MBP. Now this! A 4-bit MLX quant with 128k window should fit perfectly, in the 50-70 tok/s range.
by pwython 1mo ago
I was already rolling around the idea of a 128GB M5 Max MBP. Now this!
A 4-bit MLX quant with 128k window should fit perfectly, in the 50-70 tok/s range.
- sscaryterry 1mo agoI have a 128GB M5 Max, and it sucks at this stage. 50-70 tok/s might be something...
- smcleod 1mo ago50-70tk/s is what I get on my m5 max on a 5-6bit Qwen 3.8 27B?
- Casteil 1mo agoI don't know what black magic you're up to but I see more like 30-35t/s on a 16" M5 Max using 3.8:27b Q4, regardless of whether it's mlx or gguf. qwen3.5:122b-a10b is significantly faster at around 60-65.
- syntaxing 1mo agoWith MTP? I get 25-30 TPS on a strix halo. 50+ on a M5 max should very doable. Dflash (2) will push your TG even further
- Casteil 1mo agoIt's a bit deceptive to state inference speeds without mentioning the additional things you're doing to achieve them
- smcleod 1mo agoNo magic, just oMLX with MTP. You can look through the speed the community is getting here: https://omlx.ai/benchmarks/performance?model=qwen3.8&chip=&chip_full=M5%7CMax%7C40&quantization=&context=&pp_min=&tg_min=&sort=tg_tps&order=desc https://omlx.ai/benchmarks/performance?model=qwen3.8&chip=&c...
- sscaryterry 1mo agoI tried 8-bit, perhaps I should try 6-bit.
- irthomasthomas 1mo agoIDK, prefill speed is a bigger concern for most wokflows, like agent coding, and I heard that this is quite low on macs?
- smcleod 1mo agoThat was mainly before the M4 generation when they didn't have matmul instructions.
- jasonjmcghee 1mo agoM5 prefill is much faster than M4. I've seen benchmarks that show 4-5x faster of M5 Max vs. M4 Max. For local models you're likely using M5 Max, prefill is low thousands of tokens per second, as opposed to, say high hundreds with M4 Max. For larger dense models, some fraction of that, but similar multiple.
- smcleod 1mo agoYes, I have the M5 Max. But there was no matmul acceleration before the M4 which made things a lot slower.
- deleted 1mo ago[deleted]
- Eric_WVGG 1mo agoJust out of curiosity, why run "local-local" when you could just set up a Mini or Studio at home and query it over http? [edit] whole conversation about this in another thread https://news.ycombinator.com/item?id=49433413 https://news.ycombinator.com/item?id=49433413 I’m personally considering retiring my MBP for a Studio + 15" Air whenever this MBP ages out.
- LeBit 1mo agoThis is the way. I’m doing that. Mac Mini M4 Pro with 48G RAM as a headless llama.cpp server. I much prefer using " thin clients " as the interface to the big VMs running in my homelab
- kamranjon 1mo agoI actually do this with my MBP - it's a LLM server when I'm working - and then when I'm not it's just a really great machine for video editing and other media work.
- rdsubhas 1mo agoHow do you folks code at 40-50 tps? With an extremely lightweight harness (pi) and just 8k system and tools context, and ~40tps on qwen 3.8 27B 4-bit on low thinking mode, it still takes me nearly 30-45 mins for a basic coding session... Does it work? yeah... But I'd pick a subscription anyday...
- hgoel 1mo agoDo you find subscriptions to be meaningfully faster? I didn't really feel too much of a speed difference compared to Opus.
- latentsea 1mo agoAs someone who uses Opus daily for professional work and Qwen3.8-27B for all my private stuff, yes, Opus sub is faster for me, but I'm only rocking an R9700. If you're lucky enough to have sold a kidney on the blackmarket and purchased a 5090 and you're running ninfer, then actually... I think you'd be seeing fairly comparable performance!
- julianlam 1mo agoWhen you hear stories like "Opus 5 thought for 20 minutes and then denied my request" it really puts wind in this sails of Local LMs
- bicepjai 1mo agoThere is a finite amount of time left for these companies to become next Facebook/Google, hoarding our interaction and privacy will be a point of contention pretty soon. At that moment, Qwen will be the knight in shining armor.