3 ms·
Unlike most SWE bench submissions, Qodo Command one uses the product directly. I think that the next step is getting an official "checked" mark by the SWE benc
by itamarcode 1y ago
Unlike most SWE bench submissions, Qodo Command one uses the product directly.
I think that the next step is getting an official "checked" mark by the SWE bench team
- whymauri 1y agoI feel like the bash only SWE Bench Verified (a.k.a model + mini-swe-agent) is the closest thing to measuring the inherent ability of the model vs. the scaffolding. https://github.com/SWE-agent/mini-swe-agent https://github.com/SWE-agent/mini-swe-agent
- NitpickLawyer 1y agoThere's swe-rebench, where they take "bugs/issues" by date, and you can drag a slider on their top scores to see issues solved after the model was released (obviously only truly working for open models).