4 ms·
For reference, Mimo-v2.5-Pro scored 19% on DeepSWE 1.1. This is looking great. Fable scores 70%, Kimi K3 69%, Astra 74% (all on max effort). https://deepswe.d
by ricardobeat 18d ago
For reference, Mimo-v2.5-Pro scored 19% on DeepSWE 1.1. This is looking great.
Fable scores 70%, Kimi K3 69%, Astra 74% (all on max effort).
https://deepswe.datacurve.ai/blog/deepswe-v1-1 https://deepswe.datacurve.ai/blog/deepswe-v1-1
- Cookingboy 18d ago2.6-pro just reached 63.7% by step 10, it's on step 11 right now. Even flash reached 60.7% by step 12, and it's on step 16 now. This is so exciting lmao.
- arcanemachiner 17d agoDeepSWE is saturated now IMO, and is basically worthless. Lots of new models get around 74%. Shame too, because it was a pretty decent benchmark for a few months there.
- brookst 17d agoIt is saturated, but that doesn’t mean worthless. Seeing 72% is low-signal, but 30% is still meaningful.
- markasoftware 17d agogemini 3.8 flash is also 74% and google just started letting all their engineers use claude...go figure
- ehsankia 17d ago> and google just started letting all their engineers use claude That's misleading. 1. Having different models available is useful for A/B testing and helping improve Gemini itself. 2. They have an enterprise offering for Antigravity (their agentic coding platform), and they need to test that it works well with non-Gemini models too.
- buffalobuffalo 17d agoAlso worth taking a look at is the mimo harness. It's a fork of opencode with some new modes added for long horizon tasks. One of the better open harnesses out there at the moment.