4 ms·
I've actually been working on a solution for this problem! https://www.stet.sh/ https://www.stet.sh/ At a high level, it - Mines tasks from your merged PRs/co
by bisonbear 3mo ago
I've actually been working on a solution for this problem! https://www.stet.sh/ https://www.stet.sh/
At a high level, it
- Mines tasks from your merged PRs/commits
- Replays them in Docker containers with different harness settings (change model / reasoning effort / AGENTS.md / etc)
- Grades the patches on various attributes (tests, equivalence with human patch, code quality)
The goal is to get a sense of how agents perform on your tasks, with your context, using the tools you do.
This is currently one-shot but I'd definitely like to explore session-based benchmarks as well. There are some interesting papers that just came out on this
https://arxiv.org/abs/2606.29957 https://arxiv.org/abs/2606.29957
https://arxiv.org/abs/2606.30573 https://arxiv.org/abs/2606.30573
- achalpandey 3mo agoThank you! Both of those papers are super new and super relevant. The only big gap left is that they aren't using claude code/codex as harnesses. I'll try to reuse their constructed user sessions. PS: Your work at stet is also interesting. That's definitely a problem right now that's hard to track. The only real solution is more robust CI/CD. I have since added harder validation like essentially running a full benchmark run on every prod push.