5 ms·
We have found similar when plugging GLM 5.2 into actual benchmarks in our product. The open-source models are really dialled into the public benchmarks, until y
by lawrjone 3mo ago
We have found similar when plugging GLM 5.2 into actual benchmarks in our product. The open-source models are really dialled into the public benchmarks, until you try them in context you won't have a solid idea of how they perform (Sonnet is a higher quality model than 5.2, both in prose, reasoning, and alignment).
- flumes_whims_ 3mo agoProbably because benchmarks are leaked. Included in their training or model can cheat by finding answer online.
- ACCount37 3mo agoMost benchmarks notoriously undercount multi-turn instruction following too. And it's where open source models (and Gemini) lose big time.