2 ms·
I have a 36gb M3 Max. I tested it across quite a few different options: llama.cpp, oLMX, ollama with different options. So far ollama managed to be the most pe
by m1keil 1mo ago
I have a 36gb M3 Max.
I tested it across quite a few different options: llama.cpp, oLMX, ollama with different options.
So far ollama managed to be the most performant of them all. I will get 30 to 40 tokes/sec with it when using the -mlx version of Qwen3.8.
Whatever the sauce the ollama folks baked into the mlx + MTP mix is currently working the best out of the box.
- jantse 1mo agoThanks for sharing! Did you observe a speed difference between ollamas mlx version and the mlx-community/Qwen3.8-27B-4bit from HF ran with mlx_vlm.generate (with MTP)? Or is it the same?
- deleted 1mo ago[deleted]
- m1keil 1mo agohuh.. I'm a bit of local LLM noob so I wasn't familiar with mlx_vlm. I gave it a shot now: mlx_vlm.generate --model mlx-community/Qwen3.8-27B-4bit --prompt 'give me fizz buzz in rust' --enable-thinking --draft-kind mtp --draft-model mlx-community/Qwen3.8-27B-MTP-4bit --verbose ========== Prompt: 58 tokens, 90.717 tokens-per-sec Generation: 145 tokens, 36.392 tokens-per-sec Peak memory: 17.419 GB Speculative decoding: 2.79 accepted tokens/round (1.79 accepted drafts/round, 89.4% of drafted, avg draft 2.00) over 52 rounds Which is very close to ollama, thank you! I'm not sure if I can get rid of the drafter model, if I understand correctly, the Qwen model already includes a built in draft headers, but just having --draft-kind mtp results in about 17 t/s.
- jwr 1mo agoHmm, perhaps I should switch to an MLX version… problem is, it took quite a bit of work to get llama-server (with llama.cpp) to serve my model(s) and allow requests in non-thinking (default) and thinking modes. But 30-40 tokens/s would make a big difference.