3 ms·
I so wanted to be able to use Qwen3.8-27B locally on my fully loaded M1 Max 64GB when spider-mario recommended I use MTPLX[1], but I still found the general MLX
by bellowsgulch 16d ago
I so wanted to be able to use Qwen3.8-27B locally on my fully loaded M1 Max 64GB when spider-mario recommended I use MTPLX[1], but I still found the general MLX MTP acceleration to be so poor, I'd rather just use cheap tokens from OpenCode Zen and Go.
It's just too slow of a model. I know, I know there's new hardware, but my business paid like 4-5 grand for this MBP at the time, and I just don't feel the need to pay 7-8 grand to step up to current hardware when token spend is what it is.
[1]: https://news.ycombinator.com/item?id=49611128#49612229 https://news.ycombinator.com/item?id=49611128#49612229
- serf 16d agoslow as in the output is slow, or slow as in slow token rates? qwen3.8 has been fantastic as far as token rates are concerned for me, but the overthinking thing with higher reasoning levels takes some coercion to get right. fwiw pi and hermes both handle that model fairly well. omp required a lot of tuning. I didn't bother figuring out why, I presume it's because qwen3.8 expects reasoning declarations a bit differently. nothing a proxy can't fix.