3 ms·
I would suggest careful benchmarking. I actually tested and benchmarked, and the new Qwen3.8-27B model is actually slower with MTP on my M4 Max. MTP only gains
by jwr 2mo ago
I would suggest careful benchmarking. I actually tested and benchmarked, and the new Qwen3.8-27B model is actually slower with MTP on my M4 Max. MTP only gains anything when generating long code sequences, which is very unlikely as the model spends most of its time thinking, not generating code, even if you use it for coding (which I don't).
I get 20 tokens/s on an M4 Max (larger GPU).
- smcleod 2mo agoThat shouldn't be the case, it sounds like you've got something else going on with your setup. Here's my benchmarks: https://omlx.ai/my/fadc2127d384283f5df1fcc2c093a9f95700c6a52594bf9db837a81d3418b5ec https://omlx.ai/my/fadc2127d384283f5df1fcc2c093a9f95700c6a52... which are inline with the communities: https://omlx.ai/benchmarks/performance?sort=tg_tps&order=desc&chip_full=M5%7CMax%7C40&model=Qwen3.8-27b-awq-5 https://omlx.ai/benchmarks/performance?sort=tg_tps&order=des...