4 ms·
Recent performance data on my M1 Max 32GB MacBook using oMLX. I have been working on identifying suitable model and config for my use case and system. Using a r
by akg_67 1mo ago
Recent performance data on my M1 Max 32GB MacBook using oMLX. I have been working on identifying suitable model and config for my use case and system. Using a refactor and suggest improvements prompt for a specific Django code block using VSCode Cline extension.
---
Qwen3.8-27B-4bit, Prompt Processing (PP) 66.3 tok/s, Token Generation (TG) 11.8 tok/s
Ornith-1.5-35B-A3B-MLX-4bit, PP 379.7, TG 45.8
Ornith-1.5-35B-A3B-MLX-4bit, PP 381.5, TG 46.4
Qwen3.6-35B-A3B-mxfp4, PP 389.6, TG 47.6
Qwen3.6-35B-A3B-OptiQ-4bit, PP 342.6, TG 44.4
---
Qwen3.8-27B-4bit generally runs out of output token before completing the task though excellent partial results.
Ornith-1.5-35B-A3B-MLX-4bit seems to get in the loop often specially with tool calls.
Qwen3.6-35B-A3B-mxfp4 seems to be optimal with speed and quality output.
I am going to test Qwen3.6-35B-A3B-4bit soon with same code block just to check my intuition that any derivatives don't seem to perform better than the originals.
- madduci 1mo agoInteresting, what's your Context Window?
- akg_67 1mo agoThe above tests were done with 24k context window. Testing was mostly driven by ChatGPT analyzing oMLX server logs and suggesting changes. Finally, I settled on Qwen3.6-35B-A3B-4bit with 32,768 context window and 16,384 max tokens. --- Additional results from Qwen3.6-35B-A3B-4bit (Can't edit previous comment) Qwen3.6-35B-A3B-4bit, 329.7 PP, 41.3 TG
- visarga 1mo ago> Prompt Processing (PP) 66.3 tok/s I got 400 pp tps on a 10k token input. Your numbers seem suspiciously low, maybe the input was too short to measure properly? And this dense 27B is slow, the MoE A3B models get to 1000 tps.
- akg_67 1mo agoWhat system? If on *M1 Max 32GB* or weaker, I will be interested in learning more about your setup.