3 ms·
Supposedly this is the model card. Very impressive results. https://pbs.twimg.com/media/G6CFG6jXAAA1p0I?format=jpg&name=medium https://pbs.twimg.com/media/G6CF
by golfer 11mo ago
Supposedly this is the model card. Very impressive results.
https://pbs.twimg.com/media/G6CFG6jXAAA1p0I?format=jpg&name=medium https://pbs.twimg.com/media/G6CFG6jXAAA1p0I?format=jpg&name=...
Also, the full document:
https://archive.org/details/gemini-3-pro-model-card/page/n3/mode/2up https://archive.org/details/gemini-3-pro-model-card/page/n3/...
- tweakimp 11mo agoEvery time I see a table like this numbers go up. Can someone explain what this actually means? Is there just an improvement that some tests are solved in a better way or is this a breakthrough and this model can do something that all others can not?
- rvnx 11mo agoThis is a list of questions and answers that was created by different people. The questions AND the answers are public. If the LLM manages through reasoning OR memory to repeat back the answer then they win. The scores represent the % of correct answers they recalled.
- tylervigen 11mo agoThat is not entirely true. At least some of these tests (like HLE and ARC) take steps to keep the evaluation set private so that LLMs can’t just memorize the answers. You could question how well this works, but it’s not like the answers are just hanging out on the public internet.
- stavros 11mo agoI estimate another 7 months before models start getting 115% on Humanity's Last Exam.
- HardCodedBias 11mo agoIf you believe another thread the benchmarks are comparing Gemini-3 (probably thinking) to GPT-5.1 without thinking. The person also claims that with thinking on the gap narrows considerably. We'll probably have 3rd party benchmarks in a couple of days.
- iamdelirium 11mo agoThis is easily shown that the numbers are for GPT 5.1 thinking high. Just go to the leaderboard website and see for yourself: https://arcprize.org/leaderboard https://arcprize.org/leaderboard