3 ms·
MMLU-Pro: Gemini-3.1-Pro at 91.0 Opus-4.6 at 89.1 GPT-5.4, Kimi2.6, and DS-V4-Pro tied at 87.5 Pretty impressive
by Aliabid94 5mo ago
MMLU-Pro:
Gemini-3.1-Pro at 91.0
Opus-4.6 at 89.1
GPT-5.4, Kimi2.6, and DS-V4-Pro tied at 87.5
Pretty impressive
- ant6n 5mo agoFunny how Gemini is theoretically the best -- but in practice all the bugs in the interface mean I don't want to use it anymore. The worst is it forgets context (and lies about it), but it's very unreliable at reading pdfs (and lies about it). There's also no branch, so once the context is lost/polluted, you have to start projects over and build up the context from scratch again.
- esperent 5mo agoYeah if I could use Gemini with pi.dev that would be my choice. But Gemini CLI is just so, so bad.
- spaceman_2020 5mo agoThe sheer number of bugs and lack of meaningful improvements in Google products is a clear counterargument to the AI bull thesis If AI was so good at coding, why can’t it actually make a usable Gemini/AI Studio app?
- lazycatjumping 5mo agoI gave up on Gemini 3.1 Pro in VSCode after 2 hours. They fully refunded me.
- hodgehog11 5mo agoMost of these tests are one-prompt in nature. I've also noticed issues with the PDF reader in Gemini which was very frustrating, although it is significantly better now than it was even two weeks ago. On the contrary, now GPT-5 seems to be giving me issues. In my experience, Gemini is the most insightful model for hard problems (particularly math problems that I work on).
- Alifatisk 5mo agoYou know, with a bit of prompting, you can instruct Gemini to output the state of the conversation into a prompt that you can enter in a new chat and continue where you left off. But now with a fresh context window.
- ant6n 5mo agoNot if Gemini Lost all context already. Also, it doesn't really work well, a lot of the nuance and information simply gets lost.