3 ms·
GPT-5.6 is a really good model, and quite cheap. I can finally replace GPT-5.3-Codex for my Tool Calling in n8n. Here's my benchmark results for GPT-5.6: http
by XCSme 3mo ago
GPT-5.6 is a really good model, and quite cheap. I can finally replace GPT-5.3-Codex for my Tool Calling in n8n.
Here's my benchmark results for GPT-5.6:
https://aibenchy.com/?q=gpt-5.6 https://aibenchy.com/?q=gpt-5.6
(the high reasoning variants are still running, uploading them soon too)
EDIT: The high variants are there too, enjoy the hamsters[0].
[0]: https://aibenchy.com/showcase/?q=gpt-5.6 https://aibenchy.com/showcase/?q=gpt-5.6
- lsllc 3mo agoInteresting that Sol (low) did better than Sol (medium) in your benchmark (and is barely more expensive than Terra). I too have been using 5.3 codex as a cheap-but-good model and are switching to Terra (xhigh).
- XCSme 3mo agoHere's all 3 (medium), and GPT-5.5 It GPT-5.6 doesn't seem to be a lot smarter than 5.5, but it is faster, cheaper, more efficient and more consistent: https://aibenchy.com/compare/openai-gpt-5-6-sol-medium/openai-gpt-5-6-terra-medium/openai-gpt-5-6-luna-medium/openai-gpt-5-5-medium/ https://aibenchy.com/compare/openai-gpt-5-6-sol-medium/opena...
- ai_fry_ur_brain 3mo ago[flagged]
- paxys 3mo agoWhere's your website?
- noname120 3mo agoGiven that both Gemini 3.5 Flash (high) and Gemini 3 Flash Preview (medium) beat GPT-5.6 Sol (high) for correctness and score in your benchmarks I don’t trust them at all. The rest of the ranking also doesn’t make sense, like GPT-5.3-Codex (medium) performs better than Claude Opus 4.8 (medium) yeah sure
- XCSme 3mo agoIt's because the benchmark is not coding-only. Gemini models tend to have most knowledge for most domains, and are one of the most intelligent overall. You can check other benchmarks too, on specific categories, those models still beat other SOTA models. Regarding Opus, Anthropic models often fail to follow instructions, formatting requirements or simply refuse to answer questions (i.e. Fable). The issue with Gemini models is that they are not as good as using tools or go into weird failure modes when coding or trying to extract/generate specific data. They work amazing, until they don't...
- maxo133 3mo ago100% agree with your statement that gemini models have most knowledge in domains. They shine in so many generic topics
- XCSme 3mo agoOne example where the order seems correct, is this SVG generation test: https://aibenchy.com/showcase/?q=Gemini+3.5%2Cgpt+5.6%2C+5.3+codex%2C+claude https://aibenchy.com/showcase/?q=Gemini+3.5%2Cgpt+5.6%2C+5.3... You can see that most Gemini 3.5 generations are more correct than 5.6 Sol (the net is in the middle of the table, hamster seems reasonable and not deformed, etc.)
- desterothx 3mo agoIt doesn't seem correct at all though? the supposedly best one isn't even the best gemini flash output (the medium one looks better than the high one)