4 ms·
5.6 Terra (mid tier model) as good as Fable on DeepSWE while cheaper than Opus API pricing. Seems like a homerun.
by cbg0 3mo ago
5.6 Terra (mid tier model) as good as Fable on DeepSWE while cheaper than Opus API pricing. Seems like a homerun.
- osti 3mo agoGPT usually performs better on DeepSWE while Claude does better on FrontierCode. These two coding benchmarks are pretty much the only ones right now that's still worth taking a look at imo.
- DetroitThrow 3mo agoDeepSWE seems to strongly, strongly prefer ChatGPT models. There were also major flaws in its methodology pointed out recently, that overlap strongly with the flaws OpenAI pointed out in its SWE Verified report. I use both ChatGPT and Claude for engineering work on a daily basis, touching performance critical code to application backends to frontend work, and I've found that DeepSWE scores don't reflect my reality when I assess high quality output from the models/harnesses. Not that Opus always beats GPT 5.5., but that 5.5 is ahead of Opus on a general benchmark smells off to me.