3 ms·
This is on our roadmap! At this time we are focussing on evaluation benchmarks like SWE-bench (verified). This is a simpler benchmark and does not really map w
by kirtivr 3mo ago
This is on our roadmap!
At this time we are focussing on evaluation benchmarks like SWE-bench (verified). This is a simpler benchmark and does not really map well to investigating alerts that have a huge amount of context. But its a start.
I am wondering if we can improve upon foundation models with our reproduction <-> hypothesis loop approach.
Foundation models tend to be precision first, and context limited, so they can get sidetracked by various things.
This is an interesting problem to be working on right now!