3 ms·
Cool — but is that model really the right choice for the task? I guess it is only 3B active which helps a lot but is Gemma 4 E4B not more practical?
by dofm 2mo ago
Cool — but is that model really the right choice for the task?
I guess it is only 3B active which helps a lot but is Gemma 4 E4B not more practical?
- 0xbadcafebee 2mo agoQwen 4B and 9B should be faster and better reasoning than Gemma 4 E4B. Other good options for that much RAM are Gemma 4 12b and 31b. Gemma 4 E4B would be better for native audio, but OP is using Whisper for STT so prob doesn't matter
- dofm 2mo agoThe issue is not RAM size. It's memory bandwidth! The 12B and 31B models will be useless on a Pi 5; maybe the 12B can be persuaded to run, but it may not manage more than one token per second. It only manages 17GB/s memory bandwidth (I have seen a suggestion that the 4GB model manages more). But yes — some sort of small reasoning-oriented model (Ornith?) seems a better candidate than Qwen 35B. (Don't get me wrong, I think the 35B model is ace… just seems like at least an unusual choice here)
- petruspennanen 1mo agoHmm could try Ornith - wouldn't mind if answers came a bit faster. Probably hard to reach Qwens intelligence level though.
- aamargulies 2mo agoQwen3.5-4B would be a good (better?) candidate. It uses a gated, deltanet hybrid, so your KV cache stays nearly flat as context grows, important for RAM-constrained environments like the Pi.
- petruspennanen 2mo agoIt was chosen because it presents a new level of intelligence in this size class. The smaller gemmas abd qwens are not so smart. It was a very tight fit!