2 ms·
It does not beat Claude Sonnet 3.5 on SWE Bench (42 to Claude's 50). It chooses 4 benchmarks of the 100s of available benchmarks and then decides it "beats" Cla
by patrickhogan1 2y ago
It does not beat Claude Sonnet 3.5 on SWE Bench (42 to Claude's 50). It chooses 4 benchmarks of the 100s of available benchmarks and then decides it "beats" Claude Sonnet 3.5.
- fragmede 2y agowhat are the 100 coding benchmarks? I'm only aware of 7 and it beats Claude on 5 of them.
- patrickhogan1 2y agoI'm not aware of 100 coding benchmarks, but there are over 100 LLM benchmarks. This makes sense, as there will eventually be at least one benchmark for each human task. In addition to automated benchmarks, there are also human-rated evaluations, such as Chatbot Arena. I manually tested DeepSeek v3 against Claude 3.5 Sonnet. In my human evaluation, Claude 3.5 Sonnet outperformed DeepSeek v3, and it also outperforms DeepSeek v3 on SWE Bench. Therefore, the title of the post claiming "DeepSeek v3 beats Claude 3.5 Sonnet and is way cheaper" is wrong. That said, I was surprised by how well it performed. Its fast. Ironically, I have a paid Claude Team Plan. At the same time I was conducting the evaluations, Claude was experiencing performance issues - https://status.anthropic.com https://status.anthropic.com and DeepSeek v3 was not. This is telling for the state of chip sale restrictions.
- helloericsf 2y agoTrue. More benchmark metrics here: https://x.com/deepseek_ai/status/1872242657348710721/photo/2 https://x.com/deepseek_ai/status/1872242657348710721/photo/2