8 ms·
Understanding precisely why Gemini 3 isn't front of the pack on SWE Bench is really what I was hoping to understand here. Especially for a blog post targeted at
by aliljet 11mo ago
Understanding precisely why Gemini 3 isn't front of the pack on SWE Bench is really what I was hoping to understand here. Especially for a blog post targeted at software developers...
- cube2222 11mo agoYeah, they mention a benchmark I'm seeing the first time (Terminal-Bench 2.0) and are supposedly leading in, while for some reason SWE Bench is down from Sonnet 4.5. Curious to see some third-party testing of this model. Currently it seems to primarily improve of "general non-coding and visual reasoning" primarily, based on the benchmarks.
- nico1207 11mo agoThey are not even leading in Terminal-Bench... GPT 5.1-codex is better than Gemini 3 Pro
- pawelduda 11mo agoWhy is this particular benchmark important?
- aliljet 11mo agoThus far, this is one of the best objective evaluations of real world software engineering...
- adastra22 11mo agoIdk, Sonnet 4.5 score better than Sonnet 4.0 on that benchmark, but is markedly worse in my usage. The utility of the benchmark is fading as it is gamed.
- meowface 11mo agoI think I and many others have found Sonnet 4.5 to generally be better than Sonnet 4 for coding.
- adastra22 11mo agoMaybe if you confirm to its expectations for how you use it. 4.5 is absolutely terrible for following directions, thinks it knows better than you, and will gaslight you until specifically called out on its mistake. I have scripted prompts for long duration automated coding workflows of the fire and forget, issue description -> pull request variety. Sonnet 4 does better than you’d expect: it generates high quality mergable code about half the time. Sonnet 4.5 fails literally every time.
- pawelduda 11mo agoI'm very happy with it TBH, it has some things that annoy me a little bit: - slower compared to other models that will also do the job just fine (but excels at more complex tasks), - it's very insistent on creating loads of .MD files with overly verbose documentation on what it just did (not really what I ask it to do), - it actually deleted a file twice and went "oops, I accidentaly deleted the file, let me see if I can restore it!", I haven't seen this happen with any other agent. The task wasn't even remotely about removing anything
- adastra22 11mo agoThe last point is how it usually fails in my testing, fwiw. It usually ends up borking something up, and rather than back out and fix it, it does a 'git restore' on the file - wiping out thousands of lines of unrelated, unstaged code. It then somehow thinks it can recover this code by looking in the git history (??). And yes, I have hooks to disable 'git reset', 'git checkout', etc., and warn the model not to use these commands and why. So it writes them to a bash script and calls that to circumvent the hook, successfully shooting itself in the foot. Sonnet 4.5 will not follow directions. Because of this, you can't prevent it like you could with earlier models from doing something that destroys the worktree state. For longer-running tasks the probability of it doing this at some point approaches 100%.
- pertymcpert 11mo agoI find 4.5 a much better model FWIW.
- epolanski 11mo agoNot my experience at all, 4.5 is leagues ahead the previous models albeit not as good as Gemini 2.5.
- RamtinJ95 11mo agoI concur with the other commenters, 4.5 is a clear improvement over 4.
- svantana 11mo agoSWEBench-Verified is probably benchmaxxed at this stage. Claude isn't even the top performer, that honor goes to Doubao [1]. Also, the confidence interval for a such a small dataset is about 3 percent points, so these differences could just be up to chance. [1] https://www.swebench.com/ https://www.swebench.com/
- usaar333 11mo agoclaude 4.5 gets 82% on their own highly customized scaffolding. (parallel compute with a scoring function). That beats Doubao
- spookie 11mo agoDoes anyone trust benchmarks at this point? Genuine question. Isn't the scientific consensus that they are broken and poor evaluation tools?
- mudkipdev 11mo agoI make my own automated benchmarks
- ummonk 11mo agoIs there a tool / website that makes this process easy?
- mudkipdev 11mo agoI coded it bun and openrouter(dot)ai. I have an array of benchmarks, each benchmark has a grader (for example, checking if it equals a certain string or grade the answer automatically using another LLM). Then I save all results to a file and render the percentage correct to a graph
- energy123 11mo agoThey overly emphasize tasks with small context without noise and red herrings in the context.
- rkozik1989 11mo agoHonestly, I am inclined to think a lot of the people who are wowed by benchmarks and simple tech demos probably aren't doing very much at their day job and if they're either working on simple codebases or ones that don't have very many users(more users == more bugs found). When you throw these models at complex software projects like SOAs, big object-oriented codebases, etc. their output can be totally unusable.
- Workaccount2 11mo agoIt doesn't matter, the real benchmark is taking the community temperature on the model after a few weeks of usage.
- ramesh31 11mo ago>"It doesn't matter, the real benchmark is taking the community temperature on the model after a few weeks of usage." Indeed. It's almost impossible to truly know a model before spending a few million tokens on a real world task. It will take a step-change level advancement at this point for me to trust anything but Claude right now.
- epolanski 11mo agoImho Gemini 2.5 was by far the better model on non-trivial tasks.
- oezi 11mo agoTo this day, I still don't understand why Claude gets more acclaim for coding. Gemini 2.5 consistently outperformed Claude and ChatGPT mostly because of the much larger context.
- viraptor 11mo agoDifferent styles of usage? I see Gemini praised for being able to feed the whole project and ask changes. Which is cool and all but... I never do that. Claude for me is better for specific modifications to specific parts of the app. There's a lot of context behind what's "better".
- Libidinalecon 11mo agoI can't really explain why I have barely used Gemini. I think it was just timing with the way models came out. This will be the first time I will have a Gemini subscription and nothing else. This will be the first time I really see what it can do fully.
- 11mo ago
- ezekiel68 11mo agoI mean... it achieved 76.2% vs the leader (Claude Sonnet) at 77.2%. That's a "loss" I can deal with.