4 ms·
I was just trying that yesterday on my M4 Max with the 6bit quant. It started at 40 tps but dropped to 10 once the context loaded up.
by hypersoar 1mo ago
I was just trying that yesterday on my M4 Max with the 6bit quant. It started at 40 tps but dropped to 10 once the context loaded up.
- seanmcdirmid 1mo agofor 3.8? I can get 40 tok/sec on it (M3 Max 64GB), but I don't use it because I can get 90 tok/sec with 3.6 MoE (MTP + 6 bit quant), and I don't notice any quality improvements for my tasks using a dense model. Did you ask Gemini or DeepSeek to look at your oMLX server log to see what was going on? This can help a lot if it is just a misconfiguration.
- robotresearcher 1mo agoJust did the fun thought experiment of playing back this comment thread in my head ten years in the past and it's amusingly incomprehensible.
- dcastm 1mo agoWhat’s the context size?
- seanmcdirmid 1mo ago128k, 256k is also possible, but my tok/s drops off and the performance isn’t good.