4 ms·
Thanks for sharing! Did you observe a speed difference between ollamas mlx version and the mlx-community/Qwen3.8-27B-4bit from HF ran with mlx_vlm.generate (wit
by jantse 1mo ago
Thanks for sharing!
Did you observe a speed difference between ollamas mlx version and the mlx-community/Qwen3.8-27B-4bit from HF ran with mlx_vlm.generate (with MTP)? Or is it the same?
- deleted 1mo ago[deleted]
- m1keil 1mo agohuh.. I'm a bit of local LLM noob so I wasn't familiar with mlx_vlm. I gave it a shot now: mlx_vlm.generate --model mlx-community/Qwen3.8-27B-4bit --prompt 'give me fizz buzz in rust' --enable-thinking --draft-kind mtp --draft-model mlx-community/Qwen3.8-27B-MTP-4bit --verbose ========== Prompt: 58 tokens, 90.717 tokens-per-sec Generation: 145 tokens, 36.392 tokens-per-sec Peak memory: 17.419 GB Speculative decoding: 2.79 accepted tokens/round (1.79 accepted drafts/round, 89.4% of drafted, avg draft 2.00) over 52 rounds Which is very close to ollama, thank you! I'm not sure if I can get rid of the drafter model, if I understand correctly, the Qwen model already includes a built in draft headers, but just having --draft-kind mtp results in about 17 t/s.