4 ms·
35B MoE is certainly a good and fast local model. I find 27B dense to be quite a bit smarter, so I daily drive that. I wish there was a ~100B MoE with maybe 10B
by pettijohn 2mo ago
35B MoE is certainly a good and fast local model. I find 27B dense to be quite a bit smarter, so I daily drive that. I wish there was a ~100B MoE with maybe 10B active. It would be super smart and fast!
- nozzlegear 2mo agoI've heard 27B is smarter! I tried it some time ago but couldn't get it working with my oMLX. I need to try it again.
- mattnewton 2mo agoHonestly the 27b dense one punches way above its weight in a lot of domains, especially coding in my testing, so I think you will probably be disappointed.
- npodbielski 2mo agoIn my case I would say they are comparable but moe models are looping and getting lost a lot more than dense models. On the other hand having 90t/s with any local model is nice and Pi with loop police extension can prevent looping a lot.
- cyanydeez 2mo agoLooping seems related to quantization and not the model itself. If youre digging deep into quants to get working context then yeah.
- mattnewton 2mo agoThere was a 3.5 122B 10A release - https://huggingface.co/Qwen/Qwen3.5-122B-A10B https://huggingface.co/Qwen/Qwen3.5-122B-A10B
- kanemcgrath 2mo agoI tried it for a bit, and It was not really worth its size. It got swept up in all the other AI news recently, but laguna s 2.1 I think is the best ~100B moe model right now
- nozzlegear 2mo agoI didn't mention it above, but Laguna S is my other favorite model. I use Qwen a lot more, it's smaller and faster, but I like to switch to Laguna when I feel like I need a "heavy hitter" for certain huge or complex tasks.
- tommica 2mo agoWhat on earth hardwares do you guys have to be able to run 100gb models locally?! That's crazy! I'm here struggling to even get 27b models to run in somewhat usable way
- nozzlegear 2mo agoHaha I'm on an Mac Studio with an M1 Ultra, 64gb ram. I bought it when it first came out, it just happens to be good for local LLMs. I have to use a smaller quant of Laguna S though (I think 4-bit? Not at my machine to check), as 8-bit and full size definitely don't fit in the 64gb I have.
- lcnPylGDnU4H9OF 2mo agoYeah, a good rule of thumb is that the weights take up ~100% of the size of the model, so 100B bytes (8-bit quant) would be, well, 100GB and a 4-bit quant would be half that.
- tommica 2mo agoOh, that is a useful rule to know! Thanks!
- antman 2mo agoA 35 A3B as smart as previous gen 27B would be a sweet point
- wickedsight 2mo agoI use 27B in plan mode and 35B MoE in act mode. I noticed that is the best balance for me for consistent tool calls and intelligent planning. Takes some time to switch, but it's worth it for me.
- julianlam 2mo agoWonder if it's possible to share a common cache via cachyllama between 27B and 35B-A3B
- traceroute66 2mo ago> 35B MoE is certainly a good and fast local model. I find 27B dense to be quite a bit smarter Isn't that just the definition of MoE vs dense ?
- cyanydeez 2mo agoFull name is 35B-A3B. 3B Is the token generator thats selected out of the 35b available in the model, which is some layered jazz. So it can be dumber but its quite capablr.