6 ms·
The SWE-Bench scores are very, very high for an open source model of this size. 46.8% is better than o3-mini (with Agentless-lite) and Claude 3.6 (with AutoCode
by oofbaroomf 1y ago
The SWE-Bench scores are very, very high for an open source model of this size. 46.8% is better than o3-mini (with Agentless-lite) and Claude 3.6 (with AutoCodeRover), but it is a little lower than Claude 3.6 with Anthropic's proprietary scaffold. And considering you can run this for almost free, this is a very extraordinary model.
- falcor84 1y agoJust to confirm, are you referring to Claude 3.7?
- oofbaroomf 1y agoNo. I am referring to Claude 3.5 Sonnet New, released October 22, 2024, with model ID claude-3-5-sonnet-20241022, colloquially referred to as Claude 3.6 Sonnet because of Anthropic's confusing naming.
- SkyPuncher 1y ago> colloquially referred to as Claude 3.6 Interesting. I've never heard this.
- simonw 1y agoIt's the reason Anthropic called their next release 3.7 Sonnet - the 3.6 version number was already being used by some in the community to refer to their 3.5v2.
- turing_complete 1y agobecause nobody says that
- NiloCK 1y agoAnthropic moved from 3.5, to 3.5(new), to 3.7. They skipped 3.6 because of usage in the community, and because 3.5(newer) probably passed some threshold of awfulness. People also use 3.5.1 to refer to 3.5(new)/3.6. The remaining difficulty now is when people refer to 3.5, without specifying (new) or (old). I find most unspecified references to 3.5 these days are actually to 3.6 / 3.5.1 / 3.5(new), which is confusing.
- skerit 1y agoThat's not correct. I have always referred to it as v3.6, and I've seen plenty of other people do so too. It's why their next model was called v3.7
- Deathmax 1y agoAlso known as Claude 3.5 Sonnet V2 on AWS Bedrock and GCP Vertex AI
- ttoinou 1y agoAnd it is a very good LLM. Some people complain they don't see an improvement with Sonnet 3.7
- moffkalast 1y agoThe model formerly known as Claude 3.6 Sonnet?
- AstroBen 1y agoextraordinary.. or suspicious that the benchmarks aren't doing their job
- sagarpatil 1y agoThey are referring to SWE bench lite. Just want to make sure you are too.
- svantana 1y agoWhere did you get that idea? In the post they are repeatedly referring to SWEBench-Verified and nothing else.
- sagarpatil 1y agoSorry. I was wrong.