4 ms·
I have been contemplating a M5 Pro MBP, but for the life for me I wasn't able to find benchmarks for real-world models, do you happen to know how many tokens pe
by busfahrer 5mo ago
I have been contemplating a M5 Pro MBP, but for the life for me I wasn't able to find benchmarks for real-world models, do you happen to know how many tokens per second roughly you get with MoE models like Qwen 3.6 35B/A3B or Gemma 4 26B?
- ahknight 5mo agoI'm not normally one to share videos as answers, but this particular fellow does a LOT of work with local AIs and Macs and happens to have a nuanced answer. https://youtu.be/XGe7ldwFLSE https://youtu.be/XGe7ldwFLSE
- egorfine 5mo agoQwen 3.6 35B running on oMLX 0.3.9rc1: on oMLX I get 86 t/s on Q4 and 74 t/s on Q6. Bear in mind that ttft on MLX is much much faster on M5 Pro as compared to M4 Pro. Also bear in mind that those figures are with NO optimizations whatsoever: no MCP, no DFlash. I am waiting for both to be released for the Qwen models.
- juancn 5mo agoI'm running unsloth/Qwen3.6-35B-A3B-UD-Q8_K_XL on an M3 Max, 64GB at ~57 t/s with llama-server
- brcmthrowaway 5mo agoPrefill speed and 27B number?
- juancn 4mo agoPrefill is around ~600 t/s. I don't remember what the 27B was, I tried a 27B with different quantization at some point for that one, but I settled on the 31B.
- embedding-shape 5mo agoYou need to ask macOS people for their prefill speed as well, there are two numbers you care about here, and current MacBooks have generally terrible numbers when it comes to prefill performance. Surely it'll get better with time, but if you already have a desktop, I'd go the "beefy GPU" route first.
- egorfine 4mo ago> current MacBooks have generally terrible numbers when it comes to prefill performance Previous MacBooks. Prefill speed on M4 Pro and M5 Pro are hugely different.
- embedding-shape 4mo agoAlright, show me the numbers then, whenever I ask any macOS people about their prefill speed instead of generation speed, they all seem to disappear :P
- egorfine 4mo agohere are some: https://x.com/egorFiNE/status/2045107868261667091 https://x.com/egorFiNE/status/2045107868261667091
- egorfine 5mo agoQwen3.6 27B oQ6: 12.5 t/s generation, 340-360 t/s pp.
- egorfine 5mo agoNative MCP: For Qwen 35B enabling native MCP on MLX models slows it down by 10%. For Qwen 27B enabling native MCP on MLX models speeds token generation up almost exactly 1.5x. (all tested on M5 pro).
- mlvljr 5mo ago[dead]