4 ms·
Not OP, but I’m running local models on a M1 Max as well with 64GB RAM. It varies by model, but I’m getting 50-60 t/s with Qwen 3.6 35B and Qwen 3 coder 30B.
by victords 1mo ago
Not OP, but I’m running local models on a M1 Max as well with 64GB RAM.
It varies by model, but I’m getting 50-60 t/s with Qwen 3.6 35B and Qwen 3 coder 30B.
I’ve also used Qwen 3.8 27B but I get 10t/s on it.
It’s useable in some use cases, but I rely mostly on my $20 Claude subscription.
- copperx 1mo agoThat's so cool. I wonder if the regular M5 can run those models too.
- darthcircuit 1mo agoI run qwen 3.8 27b on my m5 mbp, with 48gb of unified ram and I’m getting around 10-15 tok/s. 3.6 35b a3b, I’m getting upwards of 100
- spider-mario 1mo agoTry 3.8 27B in MTPLX; I get about 30 tok/s with the same hardware as you. (Although it does use around 90-95W of power, compared to the ~60W that 3.6 35B-A3B uses to generate 55 tok/s. That’s about 3 J/tok instead of 1.)