4 ms·
Qwen3.8-27B runs at 59.5 tok/s on my M4 Max, 40-core GPU, 128 GB I use it occasionally for classification and other tasks but I wouldn't trust those smaller mo
by asats 1mo ago
Qwen3.8-27B runs at 59.5 tok/s on my M4 Max, 40-core GPU, 128 GB
I use it occasionally for classification and other tasks but I wouldn't trust those smaller models with the real work and for larger data processing it's too slow, e.g. a dataset I wanted to classify would've taken 56 days on my laptop vs just paying the cheap Luna prices to openai and getting it done in a few hours.
- try-working 1mo agoNot sure I would trust Luna with that. Deepseek Pro Max and Code Mode I would be more inclined to trust.
- pumanoir 1mo ago59.5 t/s is really good. Which engine/quant are you using?
- asats 1mo ago[dead]