4 ms·
huh.. I'm a bit of local LLM noob so I wasn't familiar with mlx_vlm. I gave it a shot now: mlx_vlm.generate --model mlx-community/Qwen3.8-27B-4bit --prompt 'g
by m1keil 2mo ago
huh.. I'm a bit of local LLM noob so I wasn't familiar with mlx_vlm.
I gave it a shot now:
mlx_vlm.generate --model mlx-community/Qwen3.8-27B-4bit --prompt 'give me fizz buzz in rust' --enable-thinking --draft-kind mtp --draft-model mlx-community/Qwen3.8-27B-MTP-4bit --verbose
==========
Prompt: 58 tokens, 90.717 tokens-per-sec
Generation: 145 tokens, 36.392 tokens-per-sec
Peak memory: 17.419 GB
Speculative decoding: 2.79 accepted tokens/round (1.79 accepted drafts/round, 89.4% of drafted, avg draft 2.00) over 52 rounds
Which is very close to ollama, thank you!
I'm not sure if I can get rid of the drafter model, if I understand correctly, the Qwen model already includes a built in draft headers, but just having --draft-kind mtp results in about 17 t/s.