3 ms·
Qwen 3.6 27B and 3.8 27B are the darlings of local inference at the moment. The only other thing anyone is using is Qwen 3.8 Flash Next, only by memory-rich pe
by suprjami 20d ago
Qwen 3.6 27B and 3.8 27B are the darlings of local inference at the moment.
The only other thing anyone is using is Qwen 3.8 Flash Next, only by memory-rich people.
Depending on which benchmarks you believe, these models (and the Ornith 1.5 finetune of Qwen 35B-A3B) are competitive at about Opus 4.5 to 4.7 level. That matches my experience in real tasks over the last few months.
Not bad for something you can run at home for a couple of thousand dollars.
- nasutton12 20d agothe early 1.x versions of chad were tied to ornith. i still miss the speed of that moe. https://huggingface.co/nathansutton/Ornith-1.0-35B-UD-Q2_K_XL-MLX https://huggingface.co/nathansutton/Ornith-1.0-35B-UD-Q2_K_X...
- suprjami 20d agoI skimmed chad the other day. There's very little to it (by design). afaics you should be able to replicate what chad does by copying the chad system prompt into a SYSTEM.md for Pi. Then you could use it with whatever model/provider you want.
- nasutton12 20d agopi is a fantastic harness! they are a good default in the same way llama.cpp is. it works with everything and that is the point. i was steering chad in the opposite direction. one model & one set of silicon -> taken to the max. swap out your CHAD_MODEL and it still runs, you just leave the drafter and the kernels behind.
- Terretta 18d agoLove the idea… Except the "one model" is too small for a 64GB Mac much less 128GB, sad since the Q3 is proven less competent. Offering a Q boost (with no leave behinds) on first run would be a bump worth some vibe coding while.