4 ms·
I'm encountering the same behavior. I've tried 4-8bit quants and get 14-17 tok/s with one run that achieved 19. I'm eagerly awaiting dflash2 support in Unsloth
by Atreiden 1mo ago
I'm encountering the same behavior. I've tried 4-8bit quants and get 14-17 tok/s with one run that achieved 19. I'm eagerly awaiting dflash2 support in Unsloth or LM Studio, as allegedly that should increase throughput to around 30tok/s, which is the baseline for what I consider at least somewhat interactive.
Jealous of the folks with 5090s running ninfer and getting >100tok/s. At those speeds it's a true frontier replacement IMO.
- seanmcdirmid 1mo agoThat isn’t right. Use oMLX or something similar to serve the model with internal MTP enabled. You should get at least 40 tok/sec, but I’m not sure what hardware you are using. Dflash will mess up batching, not really worth it. You need to sift through hugging face for the right model though, and some of the MTP models are meant for rapid MLX and not oMLX.
- bilbo0s 1mo agoWell, yes and no. The author has an M3. Here's reality, MLX on the software layer will not magically place hardware matrix multiplication units in your GPU cores. Newer Macs are always just gonna smoke anything earlier than an M5. Not even sure why this person is trying to get this stuff to run on hardware that wasn't designed for AI? Everyone is pointing him to newer hardware precisely because you need the newer stuff to get models to be performant. You can go with AMD, NVidia or Apple, but you're gonna be using stuff designed well after the M3 if you want to push >100tok/s.
- refulgentis 1mo ago> Not even sure why this person is trying to get this stuff to run on hardware that wasn't designed for AI? Posturing/overclaiming like this shades rather than illuminates, there are no worlds in which the "M3...wasn't designed for AI". My M4 Max 64 GB gets the same speed.
- pcf 1mo agoI would say that an M3 Ultra was designed for the LLMs around that time, plus Apple was pushing hard towards MLX over time.
- hypersoar 1mo agoI was just trying that yesterday on my M4 Max with the 6bit quant. It started at 40 tps but dropped to 10 once the context loaded up.
- seanmcdirmid 1mo agofor 3.8? I can get 40 tok/sec on it (M3 Max 64GB), but I don't use it because I can get 90 tok/sec with 3.6 MoE (MTP + 6 bit quant), and I don't notice any quality improvements for my tasks using a dense model. Did you ask Gemini or DeepSeek to look at your oMLX server log to see what was going on? This can help a lot if it is just a misconfiguration.
- robotresearcher 1mo agoJust did the fun thought experiment of playing back this comment thread in my head ten years in the past and it's amusingly incomprehensible.
- dcastm 1mo agoWhat’s the context size?
- seanmcdirmid 1mo ago128k, 256k is also possible, but my tok/s drops off and the performance isn’t good.
- anon373839 1mo agoYou and I have the same machine. Do you mind sharing the model ID you're using? I'm on oMLX also but I haven't seen anything above ~20 tok/s out of 3.8 27B, even with MTP and generating code.
- seanmcdirmid 1mo agohttps://huggingface.co/Jundot/Qwen3.6-35B-A3B-oQ6-fp16-mtp https://huggingface.co/Jundot/Qwen3.6-35B-A3B-oQ6-fp16-mtp is my favorite right now (and nothing else is even close), you should be able to get 80-90 tok/sec enabling MTP and disabling turboquant. Supposedly the 6-bit quant is essential for decent MTP performance, although I didn’t test the 4-bit quant.
- pllbnk 1mo agoI run similar workflows as the author on my 5090. It's really good and reasonably fast at ~90 tok/sec on LM Studio. I haven't tried ninfer yet. The only problem is having to be mindful about the context size. I am jealous of the folks with RTX 6000.
- wincy 1mo agoNinfer is the real deal. 160 tokens/sec. You can parallelize 4 chats at once and get ~400 tokens/sec I’ve heard. Although that just compounds the context size issue, still absolutely incredible for running locally. I made a web based battle chess game with sound effects running on a pi harness in about 15 minutes, complete with a (very unskilled) “AI” (not an llm) that plays against you if you want!
- pllbnk 1mo agoJust tried it, really cool and works as advertised.
- nialv7 1mo agodflash2 is atmost 10-20% above mtp, it won't get you to 30 tps