5 ms·
As usual, the Jinja templates are messed up so use this [0] to reduce or turn off thinking, fix tool calling, keep a 100% KV cache hit rate, etc. [0] https://h
by satvikpendem 2mo ago
As usual, the Jinja templates are messed up so use this [0] to reduce or turn off thinking, fix tool calling, keep a 100% KV cache hit rate, etc.
[0] https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates
- skrebbel 2mo agoI'd love to understand this more. Are you saying the Qwen team spends their very impressive human and compute resources on publishing these amazing models and then botches the chat template with mundane bugs? Like maybe I just misunderstand what's the hard part but wouldn't you assume that people who can put together an impressive model can also write a proper jinja chat template for it?
- Der_Einzige 2mo agoYes yes, oh god yes. They also spread FUD in the form of terrible recommended sampler settings. If you're using llamacpp, turn on top-n-sigma with sigma of 1, turn off top-p/top-k. You'll thank me later.
- skrebbel 2mo agoHow is terrible settings a case of FUD?
- MrDrMcCoy 2mo agoFor those us us who don't know, what do those parameters do and why are they better?
- deleted 2mo ago[deleted]
- suprjami 2mo agoTemperature, top-up, top-k, min-p all control which token the model predicts next and how likely it is to select one token over the other. You might understand this as "The capital of France is..." and the model isn't always going to select "Paris". Sometimes it will start a descriptive sentence or even get the answer wrong. That selection of the next token is what these settings control, and lots of sub-optimal selections compound over time to produce a junk response.
- MrDrMcCoy 2mo agoI broadly knew that about temperature, but lack the background in machine learning/statistics to differentiate top-n-sigma from top-k/top-p.
- fragmede 2mo agoSo do I, but we live in the future: https://chatgpt.com/share/6a7fc3d2-39f4-83e8-a7c6-825ddfb5e752 https://chatgpt.com/share/6a7fc3d2-39f4-83e8-a7c6-825ddfb5e7...
- suprjami 2mo agoYou might find this article relevant: https://news.ycombinator.com/item?id=49151933 https://news.ycombinator.com/item?id=49151933
- fragmede 2mo agoI've run the inference to get the answers I linked to. If someone else does the same thing, that involves extra energy. If I read your conversation instead of generating my own, then that's one less tree that has to be chopped down. Until we go advanced geothermal or we crack fusion, energy is dirty. Read my inference or link me to yours so we don't boil the planet.
- suprjami 2mo agoTop-K: example setting 20. Select only from the 20 most likely tokens. Top-P: example setting 0.9. Select tokens whose probably accumulates to this number. So say you have tokens with 0.7 then 0.2 then 0.1, the last will not be selected because the first two tokens already accumulated to >=0.9. Min-P: example setting 0.05. Don't select tokens less probable than this value. So a token with 0.1 would be considered, a token with 0.01 would not. The purpose of all of these is to exclude very unlikely next tokens.
- deleted 2mo ago[deleted]
- nullc 2mo agoDiverging from the sampler used in RL training is not good for long multi-turn results-- it's a great way to knock models into reasoning loops that wouldn't otherwise.
- Der_Einzige 2mo agoPeer reviews NeurIPS caliber paper to prove that? Because I can show you one titled "Long context generation is a sampling problem"...
- WithinReason 2mo agoShow us, we're curious. Did you upload to ArXiv yet?
- jmiskovic 2mo agoThe Qwen team published the same sampler settings for 3.8 and presumably they used those while testing on benchmark. Do you believe they could have achieved higher result with top-n-sigma?
- satvikpendem 2mo agoYes for the first question. Google of all companies didn't even get it right with Gemma for a while until recently. For some reason it doesn't seem like people can actually get these templates right.
- kzrdude 2mo agoIs the chat template used at all when they benchmark the model?
- runeblaze 2mo agothey likely use their internal infra to run benchmarks; aligning external releases with internal environments is always painful and somewhat underincentivized
- dannyw 2mo agoIt’s unlikely they are benchmarking that downstream. They probably have private benchmark scripts with purpose-specific run telemetry and logging, etc. I’m sure there is some basic testing but it might be agentic (LLM likes its own output) and maybe just some human smoke tests.
- kzrdude 2mo agoSo what's needed to solve this is that someone publicly and popularly benchmarks using the chat template. So that they have an incentive to look good using it.
- hedgehog 2mo agoYes, I was fixing issues piecemeal until I found the froggeric template, I've had to fix I think one issues with that one but it's better.
- alfiedotwtf 2mo agoThe chat templates are usually the first thing that every major release bork on, and all new model architectures end up having a ~2 week initial window of small fixes before they’re not DoA
- verdverm 2mo agoLaguna wwa standout in that they borked the quants released with the main model and had to update the next day
- suprjami 2mo agoYou have understood correctly. One really would think these companies (including Google) who spend many millions of dollars on compute could write a few hundred lines of Jinja correctly, so their investment works optimally or at all. But they don't. Then a couple of individuals on HuggingFace fix it, either a 2-person startup like Unsloth or a volunteer like froggeric. I also don't understand how this repeatedly happens.
- z4y5f3 2mo agoI did SFT / RL post-training on Qwen3 models a bit. This is an issue that dates back long ago. My favorite theory is that they had many model variants internally, each using a slightly different chat template, so when it comes to the release they are not even sure what to use any more.
- z4y5f3 2mo agoYes, and this is not the first time they messed up. They had tokenizer bugs where the trained weights do not match the template back to Qwen3 series.
- apitman 2mo agoInteresting. Why don't the unsloth guides (https://unsloth.ai/docs/models/qwen3.8 https://unsloth.ai/docs/models/qwen3.8) mention this? Do they already include the fixes in their GGUFs?
- khimaros 2mo ago@danielhanchen may be able to answer this
- satvikpendem 2mo agoThey usually do include fixes yes.
- zenoprax 2mo agoDepends on your tooling and quant? I grabbed the unsloth Q3 and it works out if the box in opencode. I had issues with OpenWebUI with a random 3.6 A3B.
- hadlock 2mo agoThanks this bumped my agent success rate from 67% to 92.5% (!!!)