3 ms·
I’ve actually been really impressed with the 27b model they recently released - amazing performance approaching 40 tok/s on m4 max and I didn’t run into any qu
by kamranjon 2mo ago
I’ve actually been really impressed with the 27b model they recently released - amazing performance approaching 40 tok/s on m4 max and I didn’t run into any quality issues in the small set of tasks I tried. Haven’t gone full coding with it yet but suspect it’s better than say a 9b or 12b model.
- embedding-shape 2mo ago> suspect it’s better than say a 9b or 12b model Whaaat, a 27b model might be better than 9b or 12b model? What would make you do such an outrageous claim?
- kamranjon 2mo agoSorry I should have clarified - I meant that a ternary 27b model would outperform a non-quantized or 8 bit quantized 9 or 12b model - which it is generally close to (or much smaller than) in size. So yeah the comparison I was trying to make was between models of equivalent size or models that could run on similarly sized hardware.
- embedding-shape 2mo agoAh yeah, that makes a ton more sense :) I mean, what you said earlier also makes sense but was too obvious, now it makes sufficient sense, thanks for explaining :)
- dofm 2mo agoI ran into some issues that are more extreme versions of Qwen's thinking loops while reasoning. It is great at the small puzzles I set for it but it did some frankly insane things on a PHP coding task I set it. It also had some issues that might be parsing/chat template stuff, tool calling oddities. I will try it again, I did try it pretty much the day it shipped and it's possible there are more improvements in their llama.cpp fork since. It would be churlish to be overcritical, mind you — the PrismML ternary stuff is an advance — but it feels like it should be applied at training. I figure we will see that, somewhere, quite soon. Did you try the BottleCap ThinkingCap Qwen post-train with the reduced thinking overhead?
- kamranjon 2mo agoI haven't spun up thinkingcap yet but I'm aware of it and am intending to try it out soon. How did you find it?
- dofm 2mo agoNot tested much but it is not noticeably worse than the underlying Qwen 3.6 27B in Q4_K_M, which in my experience is kind of a first for a Qwen fine-tune of this nature. They are almost always worse. I think it does use fewer tokens while reasoning, which is potentially useful. I need to do more testing, because any performance advantage over the 27B is useful for me on an M1 Max.