8 ms·
Benchmarks from page 4 of the model card: | Benchmark | 3 Pro | 2.5 Pro | Sonnet 4.5 | GPT-5.1 | |-----------------------|-----------
by scrlk 11mo ago
Benchmarks from page 4 of the model card:
| Benchmark | 3 Pro | 2.5 Pro | Sonnet 4.5 | GPT-5.1 |
|-----------------------|-----------|---------|------------|-----------|
| Humanity's Last Exam | 37.5% | 21.6% | 13.7% | 26.5% |
| ARC-AGI-2 | 31.1% | 4.9% | 13.6% | 17.6% |
| GPQA Diamond | 91.9% | 86.4% | 83.4% | 88.1% |
| AIME 2025 | | | | |
| (no tools) | 95.0% | 88.0% | 87.0% | 94.0% |
| (code execution) | 100% | - | 100% | - |
| MathArena Apex | 23.4% | 0.5% | 1.6% | 1.0% |
| MMMU-Pro | 81.0% | 68.0% | 68.0% | 80.8% |
| ScreenSpot-Pro | 72.7% | 11.4% | 36.2% | 3.5% |
| CharXiv Reasoning | 81.4% | 69.6% | 68.5% | 69.5% |
| OmniDocBench 1.5 | 0.115 | 0.145 | 0.145 | 0.147 |
| Video-MMMU | 87.6% | 83.6% | 77.8% | 80.4% |
| LiveCodeBench Pro | 2,439 | 1,775 | 1,418 | 2,243 |
| Terminal-Bench 2.0 | 54.2% | 32.6% | 42.8% | 47.6% |
| SWE-Bench Verified | 76.2% | 59.6% | 77.2% | 76.3% |
| t2-bench | 85.4% | 54.9% | 84.7% | 80.2% |
| Vending-Bench 2 | $5,478.16 | $573.64 | $3,838.74 | $1,473.43 |
| FACTS Benchmark Suite | 70.5% | 63.4% | 50.4% | 50.8% |
| SimpleQA Verified | 72.1% | 54.5% | 29.3% | 34.9% |
| MMLU | 91.8% | 89.5% | 89.1% | 91.0% |
| Global PIQA | 93.4% | 91.5% | 90.1% | 90.9% |
| MRCR v2 (8-needle) | | | | |
| (128k avg) | 77.0% | 58.0% | 47.1% | 61.6% |
| (1M pointwise) | 26.3% | 16.4% | n/s | n/s |
n/s = not supported
EDIT: formatting, hopefully a bit more mobile friendly
- manmal 11mo agoLooks like it will be on par with the contenders when it comes to coding. I guess improvements will be incremental from here on out.
- CjHuber 11mo agoIf it’s on par in code quality, it would be a way better model for coding because of its huge context window.
- falcor84 11mo ago> I guess improvements will be incremental from here on out. What do you mean? These coding leaderboards were at single digits about a year ago and are now in the seventies. These frontier models are arguably already better at the benchmark that any single human - it's unlikely that any particular human dev is knowledgeable to tackle the full range of diverse tasks even in the smaller SWE-Bench Verified within a reasonable time frame; to the best of my knowledge, no one has tried that. Why should we expect this to be the limit? Once the frontier labs figure out how to train these fully with self-play (which shouldn't be that hard in this domain), I don't see any clear limit to the level they can reach.
- zamadatix 11mo agoA new benchmark comes out, it's designed so nothing does well at it, the models max it out, and the cycle repeats. This could either describe massive growth of LLM coding abilities or a disconnect between what the new benchmarks are measuring & why new models are scoring well after enough time. In the former assumption there is no limit to the growth of scores... but there is also not very much actual growth (if any at all). In the latter the growth matches, but the reality of using the tools does not seem to say they've actually gotten >10x better at writing code for me in the last year. Whether an individual human could do well across all tasks in a benchmark is probably not the right question to be asking a benchmark to measure. It's quite easy to construct benchmark tasks a human can't do well in that you don't even need AI to do better.
- Alifatisk 11mo agoThese numbers are impressive, at least to say. It looks like Google has produced a beast that will raise the bar even higher. What's even more impressive is how Google came into this game late and went from producing a few flops to being the leader at this (actually, they already achieved the title with 2.5 Pro). What makes me even more curious is the following > Model dependencies: This model is not a modification or a fine-tune of a prior model So did they start from scratch with this one?
- benob 11mo agoWhat does it mean nowadays to start from scratch? At least in the open scene, most of the post-training data is generated by other LLMs.
- Alifatisk 11mo agoThey had to start with a base model, that part I am certain of
- postalcoder 11mo agoGoogle was never really late. Where people perceived Google to have dropped the ball was in its productization of AI. The Google's Bard branding stumble was so (hilariously) bad that it threw a lot of people off the scent. My hunch is that, aside from "safety" reasons, the Google Books lawsuit left some copyright wounds that Google did not want to reopen.
- Alifatisk 11mo agoOh, I remember the times when I compared Gemini with ChatGPT and Claude. Gemini was so far behind, it was barely usable. And now they are pushing the boundries.
- postalcoder 11mo agoYou could argue that chat-tuning of models falls more along the lines of product competence. I don't think there was a doubt about the upper ceiling of what people thought Google could produce.. more "when will they turn on the tap" and "can Pichai be the wartime general to lead them?"
- falcor84 11mo agoThat looks impressive, but some of the are a bit out of date. On Terminal-Bench 2 for example, the leader is currently "Codex CLI (GPT-5.1-Codex)" at 57.8%, beating this new release.
- sigmar 11mo agoThat's a different model not in the chart. They're not going to include hundreds of fine tunes in a chart like this.
- falcor84 11mo agoIt's not just one of many fine tunes; it's the default model used by OpenAI's official tools.
- Taek 11mo agoIt's also worth pointing out that comparing a fine-tune to a base model is not apples-to-apples. For example, I have to imagine that the codex finetune of 5.1 is measurably worse at non-coding tasks than the 5.1 base model. This chart (comparing base models to base models) probably gives a better idea of the total strength of each model.
- deleted 11mo ago[deleted]
- NitpickLawyer 11mo agoWhat's more impressive is that I find gemini2.5 still relevant in day-to-day usage, despite being so low on those benchmarks compared to claude 4.5 and gpt 5.1. There's something that gemini has that makes it a great model in real cases, I'd call it generalisation on its context or something. If you give it the proper context (or it digs through the files in its own agent) it comes up with great solutions. Even if their own coding thing is hit and miss sometimes. I can't wait to try 3.0, hopefully it continues this trend. Raw numbers in a table don't mean much, you can only get a true feeling once you use it on existing code, in existing projects. Anyway, the top labs keeping eachother honest is great for us, the consumers.
- HugoDias 11mo agovery impressive. I wonder if this sends a different signal to the market regarding using TPUs for training SOTA models versus Nvidia GPUs. From what we've seen, OpenAI is already renting them to diversify... Curious to see what happens next
- fariszr 11mo agoThis is a big jump in most benchmarks.And if it can match other models in coding while having that Google TPM inference speed and the actually native 1m context window, it's going to be a big hit. I hope it's isn't such a sycophant like the current gemini 2.5 models, it makes me doubt its output, which is maybe a good thing now that I think about it.
- danielbln 11mo ago> it's over for the other labs. What's with the hyperbole? It'll tighten the screws, but saying that it's "over for the other labs' might be a tad premature.
- fariszr 11mo agoI mean over in that I don't see a need to use the other models. Codex models are the best but incredibly slow. Claude models are not as good(IMO) but much faster. If gemini can beat them while having being faster and having better apps with better integrations, i don't see a reason why I would use another provider.
- nprateem 11mo agoYou should probably keep supporting competitors since if there's a monopoly/duopoly expect prices to skyrocket.
- risyachka 11mo ago> it's over for the other labs. Its not over and never will be for 2 decade old accounting software, it is definitely will not be over for other AI labs.
- xnx 11mo agoCan you explain what you mean by this? iPhone was the end of Blackberry. It seems reasonable that a smarter, cheaper, faster model would obsolete anything else. ChatGPT has some brand inertia, but not that much given it's barely 2 years old.
- Jcampuzano2 11mo agoWe knew it would be a big jump and while it certainly is in many areas - its definitely not "groundbreaking/huge leap" worthy like some were thinking from looking at these numbers. I feel like many will be pretty disappointed by their self created expectations for this model when they end up actually using it and it turns out to be fairly similar to other frontier models. Personally I'm very interested in how they end up pricing it.
- trunch 11mo agoWhich of the LiveCodeBench Pro and SWE-Bench Verified benchmarks comes closer to everyday coding assistant tasks? Because it seems to lead by a decent margin on the former and trails behind on the latter
- Snuggly73 11mo agoNeither :( LCB Pro are leet code style questions and SWE bench verified is heavily benchmaxxed very old python tasks.
- veselin 11mo agoI work a lot on testing also SWE bench verified. This benchmark in my opinion now is good to catch if you got some regression on the agent side. However, going above 75%, it is likely about the same. The remaining instances are likely underspecified despite the effort of the authors that made the benchmark "verified". From what I have seen, these are often cases where the problem statement says implement X for Y, but the agent has to simply guess whether to implement the same for other case Y' - which leads to losing or winning an instance.
- danielcampos93 11mo agoI would love to know what the increased token count is across these models for the benchmarks. I find the models continue to get better but as they do their token usage also does. Aka is model doing better or reasoning for longer?
- jstummbillig 11mo agoI think that is always something that is being worked on in parallel. Recent paradigm seems to be the models understanding when they need to use more tokens dynamically (which seems to be very much in line with how computation should generally work).
- dnw 11mo agoLooks like the best way to keep improving the models is to come up with really useful benchmarks and make them popular. ARC-AGI-2 is a big jump, I'd be curious to find out how that transfers over to everyday tasks in various fields.
- vagab0nd 11mo agoShould I assume the GPT-5.1 it is compared against is the pro version?
- spoaceman7777 11mo agoWow. They must have had some major breakthrough. Those scores are truly insane. O_O Models have begun to fairly thoroughly saturate "knowledge" and such, but there are still considerable bumps there But the _big news_, and the demonstration of their achievement here, are the incredible scores they've racked up here for what's necessary for agentic AI to become widely deployable. t2-bench. Visual comprehension. Computer use. Vending-Bench. The sorts of things that are necessary for AI to move beyond an auto-researching tool, and into the realm where it can actually handle complex tasks in the way that businesses need in order to reap rewards from deploying AI tech. Will be very interesting to see what papers are published as a result of this, as they have _clearly_ tapped into some new avenues for training models. And here I was, all wowed, after playing with Grok 4.1 for the past few hours! xD
- rvnx 11mo agoThe problem is that we know in advance what is the benchmark, so Humanity's Last Exam for example, it's way easier to optimize your model when you have seen the questions before.
- stego-tech 11mo agoThis. A lot of boosters point to benchmarks as justification of their claims, but any gamer who spent time in the benchmark trenches will know full well that vendors game known tests for better scores, and that said scores aren’t necessarily indicative of superior performance. There’s not a doubt in my mind that AI companies are doing the same.
- Feuilles_Mortes 11mo agoshouldn't we expect that all of the companies are doing this optimization, though? so, back to level playing field.
- eldenring 11mo agoIts the other way around too, HLE questions were selected adversarially to reduce the scores. I'd guess even if the questions were never released, and new training data was introduced, the scores would improve.
- deleted 11mo ago[deleted]
- scrollop 11mo agoUsed an AI to populate some of 5.1 thinking's results. Benchmark | Gemini 3 Pro | Gemini 2.5 Pro | Claude Sonnet 4.5 | GPT-5.1 | GPT-5.1 Thinking ---------------------------|--------------|----------------|-------------------|---------|------------------ Humanity's Last Exam | 37.5% | 21.6% | 13.7% | 26.5% | 52% ARC-AGI-2 | 31.1% | 4.9% | 13.6% | 17.6% | 28% GPQA Diamond | 91.9% | 86.4% | 83.4% | 88.1% | 61% AIM 2025 | 95.0% | 88.0% | 87.0% | 94.0% | 48% MathArena Apex | 23.4% | 0.5% | 1.6% | 1.0% | 82% MMMU-Pro | 81.0% | 68.0% | 68.0% | 80.8% | 76% ScreenSpot-Pro | 72.7% | 11.4% | 36.2% | 3.5% | 55% CharXiv Reasoning | 81.4% | 69.6% | 68.5% | 69.5% | N/A OmniDocBench 1.5 | 0.115 | 0.145 | 0.145 | 0.147 | N/A Video-MMMU | 87.6% | 83.6% | 77.8% | 80.4% | N/A LiveCodeBench Pro | 2,439 | 1,775 | 1,418 | 2,243 | N/A Terminal-Bench 2.0 | 54.2% | 32.6% | 42.8% | 47.6% | N/A SWE-Bench Verified | 76.2% | 59.6% | 77.2% | 76.3% | N/A t2-bench | 85.4% | 54.9% | 84.7% | 80.2% | N/A Vending-Bench 2 | $5,478.16 | $573.64 | $3,838.74 | $1,473.43| N/A FACTS Benchmark Suite | 70.5% | 63.4% | 50.4% | 50.8% | N/A SimpleQA Verified | 72.1% | 54.5% | 29.3% | 34.9% | N/A MMLU | 91.8% | 89.5% | 89.1% | 91.0% | N/A Global PIQA | 93.4% | 91.5% | 90.1% | 90.9% | N/A MRCR v2 (8-needle) | 77.0% | 58.0% | 47.1% | 61.6% | N/A Argh it doesn't come out write in HN
- scrollop 11mo agoUsed an AI to populate some of 5.1 thinking's results. Benchmark..................Description...................Gemini 3 Pro....GPT-5.1 (Thinking)....Notes Humanity's Last Exam.......Academic reasoning.............37.5%..........52%....................GPT-5.1 shows 7% gain over GPT-5's 45% ARC-AGI-2...................Visual abstraction.............31.1%..........28%....................GPT-5.1 multimodal improves grid reasoning GPQA Diamond................PhD-tier Q&A...................91.9%..........61%....................GPT-5.1 strong in physics (72%) AIME 2025....................Olympiad math..................95.0%..........48%....................GPT-5.1 solves 7/15 proofs correctly MathArena Apex..............Competition math...............23.4%..........82%....................GPT-5.1 handles 90% advanced calculus MMMU-Pro....................Multimodal reasoning...........81.0%..........76%....................GPT-5.1 excels visual math (85%) ScreenSpot-Pro..............UI understanding...............72.7%..........55%....................Element detection 70%, navigation 40% CharXiv Reasoning...........Chart analysis.................81.4%..........69.5%.................N/A
- roman_soldier 11mo agoWhy is Grok 4.1 not in the benchmarks?
- HardCodedBias 11mo agoBig if true. I'll wait for the official blog with benchmark results. I suspect that our ability to benchmark models is waning. Much more investment required in this area, but what is the play out?
- visioninmyblood 11mo agoreally great results although the results are so high i was trying a simple example of object detection and the performance was kind of poor in agentic frameworks. Need to see how this performs on other other tasks.
- keyle 11mo agoThe vending-bench 2 benchmark is kind of nutty [1]. Not sure 360 days is enough of a sample really but it's an interesting take on AI benchmarks. Are there any other interesting benchmarks to look at? [1] https://andonlabs.com/evals/vending-bench-2 https://andonlabs.com/evals/vending-bench-2
- AISnakeOil 11mo agonice numbers, but what does this actually mean? What does this model do that others can't already.
- spwa4 11mo agoBut ... what's missing from this comparison: Kimi-K2. When ChatGPT-3 exploded, OpenAI had at least double the benchmark scores of any other model, open or closed. Gemini 3 Pro (not the model they actually serve) outperforms the best open model ... wait it does not uniformly beat the best open model anymore. Not even close. Kimi-k2 beats Gemini 3 pro on several benchmarks. On average it scores just under 10% better then the best open model, currently Kimi-K2. Gemini-3 pro is in fact only the best in about half the benchmarks tested there. In fact ... this could be another llama4 moment. The reason Gemini-3 pro is the best model is a very high score on a single benchmark ("Humanity's last exam"), if you take that benchmark out GPT-5.1 remains the best model available. The other big improvement is "SciCode", and if you take that out too the best open model, Kimi K2, beats Gemini 3 pro. https://artificialanalysis.ai/models https://artificialanalysis.ai/models And then, there's the pricing: Kimi K2 on OpenRouter: $0.50 / M input tokens, $2.40 / M output tokens Gemini 3 Pro: For contexts ≤ 200,000 tokens: US$ 2.00 per 1 M input tokens, Output tokens: US$ 12.00 per 1 M tokens For contexts > 200,000 tokens (long context tier): US$ 4.00 per 1 M input tokens , US$ 18.00 per 1 M output tokens So Gemini 3 pro is 4 times, 400%, the price of the best open model (and just under 8 times, 800%, with long context), and 70% more expensive than GPT-5.1 The closed models in general, and Google specifically, serve Gemini 3 pro at double to triple the speed (as in tokens-per-second) of openrouter. Although even here it is not the best, that's openrouter with gpt-oss-120b.