3 ms·
Google provided incorrect settings and an imperfect template. Unsloth modified the template and then finetuned their own version of the model to optimize for s
by CMay 2mo ago
Google provided incorrect settings and an imperfect template.
Unsloth modified the template and then finetuned their own version of the model to optimize for some benchmarks as a means of validating quants.
Google and Llama.cpp then adopt template changes by default, so anyone downloading the new model or even using the original model will now automatically be using it incorrectly.
Llama.cpp also uses the same inference setting defaults regardless which version of the model you use and some settings are simply defaults it uses for all models.
Then even if you account for all of these, you have to be using Gemma 4 itself correctly, which many people do not.
All of these little changes and inconsistencies hurt some of the model's original capabilities. Even if you go directly to Google's repo and download the full float 16 weights with the template they have there now, you cannot simply assume you're getting the best results.
- mlvljr 2mo ago[dead]
- DanielHB 2mo agoI am very much a beginner to local LLM stuff and I find it incredibly hard to figure out how to run models optimally with the correct settings for my hardware. The number of different variations of the same model and how each quant work is super confusing as well. When I tried to run llama.cpp directly I was getting max 9tk/s on qwen3.5-9B, then I tried LM Studio with the same model and got 77tk/s. I haven't figured out yet how to get MTP working properly in either.
- car 2mo agoIf you are on Mac, have a look at the Llama-macOS app. They claim sensible settings for the linked model downloads. I'd expect the authors of Llama.cpp and the Huggingface folks to know this stuff. https://news.ycombinator.com/item?id=49328008 https://news.ycombinator.com/item?id=49328008
- tiahura 2mo agoIf your package manager / configurator isn’t claude code or codex, you’re wasting time.
- dofm 2mo agoUnless, of course, your goal is to actually understand what is going on, regardless of whether that is difficult.
- blactuary 2mo agoSome of us prefer to avoid Anthropic/OpenAI
- throw1234567891 2mo agoYour funny.
- roosterIllusi0n 2mo agoI have had luck telling the free chatgpt my graphics card brand/vram and asking it to recommend latest qwen3.8 or gemma4 model variants. Then I pasted in my server command and ask it to optimize it. I also pasted in the token per second logs to get further tweaks. If you paste the token per second info log info back to chatgpt, you can iterate with the free chatgpt to get better settings.
- agile-gift0262 2mo agoAnd what's the right way to use Gemma? Where can I find the correct template and settings if those aren't the ones provided by Google, Unsloth, and aren't built into llama.cpp? I discarded using Gemma 4 because it got into weird loops when tool calling
- kzrdude 2mo agoSome weeks ago a new official Gemma 4 release was posted that corrected some of the chat template problems. So the official release files on hugging face should be the way to go.
- CMay 2mo agoThe updated version will handle tool calling better by default, but the reasoning quality is no longer preserved and is mutilated quite badly.
- rpbiwer2 2mo agoSo then how do you run it unmutilated?
- hedgehog 2mo agoLog all the calls and run on a periodic cadence (cron or ever N turns) a larger model (like Opus) to read samples of the traces and edit the template to fix observed problems. There are some signs that help find interesting things to look at, errors of course, but also overly long responses, prefix cache misses, tool call errors, etc.
- CMay 2mo agoDownload the original model with the original template, not updated versions of the model or finetuned versions of the model. Then create your own reasoning tests to verify that it is working correctly. You can set a specific seed value to make sure the generation is the same every time, that way you can identify any tokens that are different. Afterwards, try making small incremental changes to the template and validate your tests each time in order to try to adopt the improvements from the newer templates. If the reasoning quality degrades, undo your changes and try again or test alternative solutions.
- rao-v 2mo agoI run llama.cpp and specialized forks on 64GB of HBM and I still cannot figure out where to find the final correct guidance on using the Gemma 4 models. Would appreciate any kind of pointer to the latest!