3 ms·
Seems threads about local LLMs on Apple hardware feature comments listing M3/4/5 at 48GB 64GB and not 128GB. That is, users with M-series hardware that have le
by mistersquid 1mo ago
Seems threads about local LLMs on Apple hardware feature comments listing M3/4/5 at 48GB 64GB and not 128GB.
That is, users with M-series hardware that have less-than-max RAM share results whereas users with max RAM do not.
Speculating (not extrapolating), maybe users with machine that have max RAM are less interested in running local LLMs and are less averse to paying services for compute?
Personally, I’d love to see what output max RAM M-series Apple hardware in these threads.
- asats 1mo agoQwen3.8-27B runs at 59.5 tok/s on my M4 Max, 40-core GPU, 128 GB I use it occasionally for classification and other tasks but I wouldn't trust those smaller models with the real work and for larger data processing it's too slow, e.g. a dataset I wanted to classify would've taken 56 days on my laptop vs just paying the cheap Luna prices to openai and getting it done in a few hours.
- try-working 1mo agoNot sure I would trust Luna with that. Deepseek Pro Max and Code Mode I would be more inclined to trust.
- pumanoir 1mo ago59.5 t/s is really good. Which engine/quant are you using?
- asats 1mo ago[dead]
- digikata 1mo agoI've been using Qwen3.6-37B-A3B on an M1 Max w/ llama.cpp and for my practical uses I prefer it to qwen3.8. When 3.8 does answer its slower and, qualitatively, marginally better than qwen3.6, but 3.8 often ends up in unresolved thought loops and runs slower. The Moe 3.6 on my setup is much faster, 500t/s peaks, 30t/s typical, vs 3.8 150 peak, 4-9 t/s typical. While I've spend a little time tuning, I'm assuming there will be deeper tuning for 3.8 that might close the gap.