3 ms·
I'm not following the local mdoel scene too closely but this seems quite amazing. Is this able to be run on Apple silicon too?
by abraxas 15d ago
I'm not following the local mdoel scene too closely but this seems quite amazing. Is this able to be run on Apple silicon too?
- kamranjon 15d ago"Ternary Bonsai 2 27B reaches up to 143 tokens/second on NVIDIA GeForce RTX 5090 and 46.8 tokens/second on M5 Max. On an RTX 4090, Ternary Bonsai 2 27B consumes just 0.714 mWh/token, making it 40% more energy-efficient than an 8B model running in full-precision."
- pizza234 15d agoTheir mention of the 5090 is bit odd, since on 32 GB GPUs, Q6 fits while having better quality. Very interesting model for 16 GB GPUs though!
- sisve 15d agoThey mention 5090 with regards to speed, Q6 will not have that speed? And speed matters a lot for many use cases
- selectodude 15d ago150 tokens per second on a ternary model implies that it’s GPU bound, I’d bet a Q6 model is even faster because it’s existed longer and seen more optimization. You’d have to be insane to not run an NVFP4 quant over a ternary quant on Blackwell if they both fit.
- wincy 15d agoWith Ninfer and Qwen 3.8 27b it uses a groupwise int mixed quant, and it gets 160 tokens/sec. The mixed quant is between 4 and 6 bits.
- Foobar8568 15d agoNinfer is compatible with a nvfp4 model for the 27b. Also nowadays I prefer to use the byteshape one, I get less loops, and I am not sure if I really see a difference in speed or quality. Pure vibe agentic coding on a C++ codebase or ocaml one, ocaml one has codex as reviewer as I am more interested by that project, the other is more for fun.
- pwython 15d agoSometimes you want a decent model running in the background that doesn't take up all the VRAM.
- blurbleblurble 15d agoOr maybe even to run parallel threads of the same model!
- azatom 15d agoit is like "my fridge is 2mkm (millikilometer) from my desk" m=0.001 h=3600 it should be just Ws or just J
- Havoc 15d agoTheir first 27B bonsai was able to run on an iphone.