4 ms·
Yes for the first question. Google of all companies didn't even get it right with Gemma for a while until recently. For some reason it doesn't seem like people
by satvikpendem 2mo ago
Yes for the first question. Google of all companies didn't even get it right with Gemma for a while until recently. For some reason it doesn't seem like people can actually get these templates right.
- kzrdude 2mo agoIs the chat template used at all when they benchmark the model?
- runeblaze 2mo agothey likely use their internal infra to run benchmarks; aligning external releases with internal environments is always painful and somewhat underincentivized
- dannyw 2mo agoIt’s unlikely they are benchmarking that downstream. They probably have private benchmark scripts with purpose-specific run telemetry and logging, etc. I’m sure there is some basic testing but it might be agentic (LLM likes its own output) and maybe just some human smoke tests.
- kzrdude 2mo agoSo what's needed to solve this is that someone publicly and popularly benchmarks using the chat template. So that they have an incentive to look good using it.