3 ms·
I gave the same prompt (a small rust project that's not easy, but not overly sophisticated) to both Gemma-4 26b and Qwen 3.5 27b via OpenCode. Qwen 3.5 ran for
by swalsh 6mo ago
I gave the same prompt (a small rust project that's not easy, but not overly sophisticated) to both Gemma-4 26b and Qwen 3.5 27b via OpenCode. Qwen 3.5 ran for a bit over an hour before I killed it, Gemma 4 ran for about 20 minutes before it gave up. Lots of failed tool calls.
I asked codex to write a summary about both code bases.
"Dev 1" Qwen 3.5
"Dev 2" Gemma 4
Dev 1 is the stronger engineer overall. They showed better architectural judgment, stronger completeness, and better maintainability instincts. The weakness is execution rigor: they built more, but didn’t verify enough, so important parts don’t actually hold up cleanly.
Dev 2 looks more like an early-stage prototyper. The strength is speed to a rough first pass, but the implementation is much less complete, less polished, and less dependable. The main weakness is lack of finish and technical rigor.
If I were choosing between them as developers, I’d take Dev 1 without much hesitation.
Looking at the code myself, i'd agree with codex.
- coder543 6mo agoThere are issues with the chat template right now[0], so tool calling does not work reliably[1]. Every time people try to rush to judge open models on launch day... it never goes well. There are ~always bugs on launch day. [0]: https://github.com/ggml-org/llama.cpp/pull/21326 https://github.com/ggml-org/llama.cpp/pull/21326 [1]: https://github.com/ggml-org/llama.cpp/issues/21316 https://github.com/ggml-org/llama.cpp/issues/21316
- emidoots 6mo agowas just merged
- coder543 6mo agoIt was just an example of a bug, not that it was the only bug. I’ve personally reported at least one other for Gemma 4 on llama.cpp already. In a few days, I imagine that Gemma 4 support should be in better shape.
- stavros 6mo agoWhat causes these? Given how simple the LLM interface is (just completion), why don't teams make a simple, standardized template available with their model release so the inference engine can just read it and work properly? Can someone explain the difficulty with that?
- Yukonv 6mo agoThe model does have the format specified but there is no _one_ standard. For this model it’s defined in the [ tokenizer_config.json [0]. As for llama.cpp they seem to be using a more type safe approach to reading the arguments. [0] https://huggingface.co/google/gemma-4-31B-it/blob/main/tokenizer_config.json#L37 https://huggingface.co/google/gemma-4-31B-it/blob/main/token...
- stavros 6mo agoHm, but surely there will be converters for such simple formats? I'm confused as to how there can be calling bugs when the model already includes the template.
- zozbot234 6mo agoThe models are not technically comparable: the Qwen is dense, the Gemma is MoE. The ~33B models are the other way around!
- petu 6mo agoQwen 3.5 27B is dense, so (I think) should be compared to Gemma 4 31B. Or Gemma-4 26B(-A4B) should be compared to Qwen 3.5 35B(-A3B)
- redman25 6mo agoExactly, compare MoE with MoE and dense with dense otherwise it's apples and oranges.
- swalsh 6mo agoIts coding to coding. I could care less how the model is architected, i only care how it performs in a real world scenario.
- daemonologist 6mo agoThe implication is that there is (should be) a major speed difference - naively you'd expect the MoE to be 10x faster and cheaper, which can be pretty relevant on real world tasks.
- petu 6mo agoIf you don't care about how it's architectured, why you care about size? Compare it to Q3.5 397B-A17B. Just like smaller size models are speed / cost optimization, so is MoE. G4 26B-A4B goes 150 t/s on 4090/5090, 80 t/s on M5 Max. Q3.5 35B-A3B is comparably fast. They are flash-lite/nano class models. G4 31B despite small increase in total parameter count is over 5 times slower. Q3.5 27B is comparably slow. They are approximating flash/mini class models (I believe sizes of proprietary models in this class are closer to Q3.5 122B-A10B or Llama 4 Scout 109B-A17B).