2 ms·
It was our original assumption. Yet, we went through trajectories and agents did not recall solution. It is with a sharp contrast with task for which agents mag
by stared 3mo ago
It was our original assumption. Yet, we went through trajectories and agents did not recall solution. It is with a sharp contrast with task for which agents magically generate solution, e.g. https://openai.com/index/why-we-no-longer-evaluate-swe-bench-verified/ https://openai.com/index/why-we-no-longer-evaluate-swe-bench....
In a few instances (we covered it in Caveats) Gemini 3.5 Flash "knew" which level it was, but misremembered, and went with a wrong solution.