3 ms·
You're wrong in lots of ways. Some model cards do show regressions on benchmarks for newer models on specific tasks: https://storage.googleapis.com/deepmind-me
by suttontom 4mo ago
You're wrong in lots of ways.
Some model cards do show regressions on benchmarks for newer models on specific tasks: https://storage.googleapis.com/deepmind-media/Model-Cards/Gemini-3-1-Pro-Model-Card.pdf https://storage.googleapis.com/deepmind-media/Model-Cards/Ge...
This wasn't a new model but updates to models backed by numbers being better can make the model worse: https://openai.com/index/sycophancy-in-gpt-4o/ https://openai.com/index/sycophancy-in-gpt-4o/
The slight increases in performance/benchmarks may be just noise: https://arxiv.org/pdf/2602.07150 https://arxiv.org/pdf/2602.07150