6 ms·
Very impressive! But under arguably the most important benchmark -- SWE-bench verified for real-world coding tasks -- Claude 3.7 still remains the champion.[1]
by georgewsinger 1y ago
Very impressive! But under arguably the most important benchmark -- SWE-bench verified for real-world coding tasks -- Claude 3.7 still remains the champion.[1]
Incredible how resilient Claude models have been for best-in-coding class.
[1] But by only about 1%, and inclusive of Claude's "custom scaffold" augmentation (which in practice I assume almost no one uses?). The new OpenAI models might still be effectively best in class now (or likely beating Claude with similar augmentation?).
- oofbaroomf 1y agoClaude got 63.2% according to the swebench.com leaderboard (listed as "Tools + Claude 3.7 Sonnet (2025-02-24)).[0] OpenAI said they got 69.1% in their blog post. [0] swebench.com/#verified
- awestroke 1y agoOpenAI have not shown themselves to be trustworthy, I'd take their claims with a few solar masses of salt
- georgewsinger 1y agoYes, however Claude advertised 70.3%[1] on SWE bench verified when using the following scaffolding: > For Claude 3.7 Sonnet and Claude 3.5 Sonnet (new), we use a much simpler approach with minimal scaffolding, where the model decides which commands to run and files to edit in a single session. Our main “no extended thinking” pass@1 result simply equips the model with the two tools described here—a bash tool, and a file editing tool that operates via string replacements—as well as the “planning tool” mentioned above in our TAU-bench results. Arguably this shouldn't be counted though? [1] https://www.anthropic.com/_next/image?url=https%3A%2F%2Fwww-cdn.anthropic.com%2Fimages%2F4zrzovbb%2Fwebsite%2F08bba4487fb5ac1ba52540ee656d7e4da10ca1be-1920x1145.png&w=1920&q=75 https://www.anthropic.com/_next/image?url=https%3A%2F%2Fwww-...
- tedsanders 1y agoI think you may have misread the footnote. That simpler setup results in the 62.3%/63.7% score. The 70.3% score results from a high-compute parallel setup with rejection sampling and ranking: > For our “high compute” number we adopt additional complexity and parallel test-time compute as follows: > We sample multiple parallel attempts with the scaffold above > We discard patches that break the visible regression tests in the repository, similar to the rejection sampling approach adopted by Agentless; note no hidden test information is used. > We then rank the remaining attempts with a scoring model similar to our results on GPQA and AIME described in our research post and choose the best one for the submission. > This results in a score of 70.3% on the subset of n=489 verified tasks which work on our infrastructure. Without this scaffold, Claude 3.7 Sonnet achieves 63.7% on SWE-bench Verified using this same subset.
- georgewsinger 1y agoSomehow completely missed that, thanks! I think reading this makes it even clearer that the 70.3% score should just be discarded from the benchmarks. "I got a 7%-8% higher SWE benchmark score by doing a bunch of extra work and sampling a ton of answers" is not something a typical user is going to have already set up when logging onto Claude and asking it a SWE style question. Personally, it seems like an illegitimate way to juice the numbers to me (though Claude was transparent with what they did so it's all good, and it's not uninteresting to know you can boost your score by 8% with the right tooling).
- ianbutler 1y agoIt isn't on the benchmark https://www.swebench.com/#verified https://www.swebench.com/#verified The one on the official leaderboard is the 63% score. Presumably because of all the extra work they had to do for the 70% score.
- swyx 1y agothey also gave more detail on their SWEBench scaffolding here https://www.latent.space/p/claude-sonnet https://www.latent.space/p/claude-sonnet
- jjani 1y agoGemini 2.5 Pro is widely considered superior to 3.7 Sonnet now by heavy users, but they don't have an SWE-bench score. Shows that looking at one such benchmark isn't very telling. Main advantage over Sonnet being that it's better at using a large amount of context, which is enormously helpful during coding tasks. Sonnet is still an incredibly impressive model as it held the crown for 6 months, which may as well be a decade with the current pace of LLM improvement.
- unsupp0rted 1y agoMain advantage over Sonnet is Gemini 2.5 doesn't try to make a bunch of unrelated changes like it's rewriting my project from scratch.
- itsmevictor 1y agoI find Gemini 2.5 truly remarkable and overall better than Claude, which I was a big fan of
- enraged_camel 1y agoStill doesn't work well in Cursor unfortunately.
- ai-christianson 1y agoWorks well in RA.Aid --in fact I'd recommend it as the default model in terms of overall cost and capability.
- plantain 1y agoWorking fine here. What problems do you see?
- michaelbarton 1y agoNot the OP but believe they could be referring to the fact it’s not supported in edit mode yet, only agent mode. So far for me that’s not been too much of a roadblock. Though I still find overall Gemini struggles with more obscure issues such as SQL errors in dbt
- lattalayta 1y agoI haven't been following them that closely, but are people finding these benchmarks relevant? It seems like these companies could just tune their models to do well on particular benchmarks
- emp17344 1y agoThat’s exactly what’s happening. I’m not convinced there’s any real progress occurring here.
- mickael-kerjean 1y agoThe benchmark is something you can optimize for, doesn't mean it generalize well. Yesterday I tried for 2 hours to get claude to create a program that would extract data from a weird adobe file. 10$ later, the best I had is a program that was doing something like: switch(testFile) { case "test1.ase": // run this because it's a particular case case "test2.ase": // run this because it's a particular case default: // run something that's not working but that's ok because the previous case should // give the right output for all the test files ... }
- thefourthchime 1y agoAlso, if you're using Cursor AI, it seems to have much better integration with Claude where it can reflect on its own things and go off and run commands. I don't see it doing that with Gemini or the O1 models.
- pizzathyme 1y agoThe image generation improvement with o4-mini is incredible. Testing it out today, this is a step change in editing specificity even from the ChatGPT 4o LLM image integration just a few weeks ago (which was already a step change). I'm able to ask for surgical edits, and they are done correctly. There isn't a numerical benchmark for this that people seem to be tracking but this opens up production-ready image use cases. This was worth a new release.
- mchusma 1y agoThanks for sharing that. that was more interesting then their demo. I tried it and it was pretty good! I have felt that the ability to iterate from images blocked this from any real production use I had. This may be good enough now. Example of edits (not quite surgical but good): https://chatgpt.com/share/68001b02-9b4c-8012-a339-73525b824666 https://chatgpt.com/share/68001b02-9b4c-8012-a339-73525b8246...
- ec109685 1y agoI don’t know if they let you share the actual images when sharing a chat. For me, they are blank.
- ilaksh 1y agowait, o4-mini outputs images? What I thought I saw was the ability to do a tool call to zoom in on an image. Are you sure that's not 4o?
- knes 1y agoRight now the Swe-Bench leader Augment Agent still use Claude 3.7 in combo with o1. https://www.augmentcode.com/blog/1-open-source-agent-on-swe-bench-verified-by-combining-claude-3-7-and-o1 https://www.augmentcode.com/blog/1-open-source-agent-on-swe-... The findings are open sourced on a repo too https://github.com/augmentcode/augment-swebench-agent https://github.com/augmentcode/augment-swebench-agent
- ksec 1y agoI often wonder if we could expect that to reach 80% - 90% within next 5 years.