3 ms·
They used to compare to competing models from Anthropic, Google DeepMind, DeepSeek, etc. Seems that now they only compare to their own models. Does this mean th
by minadotcom 10mo ago
They used to compare to competing models from Anthropic, Google DeepMind, DeepSeek, etc. Seems that now they only compare to their own models. Does this mean that the GPT-series is performing worse than its competitors (given the "code red" at OpenAI)?
- poormathskills 10mo agoOpenAI has never compared their models to models from other labs in their blog post. Open literally any past model launch post to see that.
- tabletcorry 10mo agoThe matrix required for a fair comparison is getting too complicated, since you have to compare chat/thinking/pro against an array of Anthropic and Google models. But they publish all the same numbers, so you can make the full comparison yourself, if you want to.
- Tiberium 10mo agoThey did compare it to other models: https://x.com/OpenAI/status/1999182104362668275 https://x.com/OpenAI/status/1999182104362668275 https://i.imgur.com/e0iB8KC.png https://i.imgur.com/e0iB8KC.png
- enlyth 10mo agoThis looks cherry-picked, for example Claude Opus had a higher score on SWE-Bench Verified so they conveniently left it out, also GDPval is literally a benchmark made by OpenAI
- minadotcom 10mo agoagreed.
- tobias2014 10mo agoAnd who believes that the difference between 91.9% and 92.4% is significant in these benchmarks? Clearly these have margins of error that are swept under the rug.
- whimsicalism 10mo agouh oh, where did SWE bench go :D
- whimsicalism 10mo agomaybe they will release with gpt-5.2-codex
- sergdigon 10mo agoThe fact that the post is comparing their reasoning model against gemini 3 pro (the "non reasoning" model) and not gemini 3 pro deep think (the reasoning one) is quite nasty. If you compare GPT5.2 thinking to gemini 3 pro deep think, the scores are quite similar (sometimes one is better sometimes the other one is)
- Workaccount2 10mo agoThey are taking a page out of Apple's book. Apple only compares to themselves. They don't even acknowledge the existence of others.